To be honest, I just don’t find this problem space that interesting anymore. At AI cluster scale, accelerator time is vastly more expensive than redundant network capacity. It’s going to be this way for at least the next half decade. And once you control the endpoints, NICs, topology, and workload, “just massively overprovision the fabric to prevent congestion entirely” starts looking like a pretty compelling congestion-control algorithm. Once that half decade is up and these inefficiencies become a bottleneck again, AI is going to be able to vibecode a slopgorithm on NS3 that outperforms all of these and makes it moot.
It’s just not the right time to solve this problem anymore.
To be honest, I just don’t find this problem space that interesting anymore. At AI cluster scale, accelerator time is vastly more expensive than redundant network capacity. It’s going to be this way for at least the next half decade. And once you control the endpoints, NICs, topology, and workload, “just massively overprovision the fabric to prevent congestion entirely” starts looking like a pretty compelling congestion-control algorithm. Once that half decade is up and these inefficiencies become a bottleneck again, AI is going to be able to vibecode a slopgorithm on NS3 that outperforms all of these and makes it moot.
It’s just not the right time to solve this problem anymore.
https://lwn.net/Articles/1003059/
https://www.usenix.org/system/files/atc21-ousterhout.pdf
Thanks, those are great! Added to toptext as well.