In the AI infrastructure race, achieving peak performance promised by the specs of a chip is very much gated on the fidelity of the software tooling stack associated with it.
Over the past several years, the semiconductor industry has poured hundreds of billions of dollars into alternative accelerators, hyperscaler chips, and specialized edge hardware. Yet, despite massive advances in raw performance specifications, the majority of those in the AI silicon ecosystem struggle to achieve their theoretical capacity when running production workloads.
Reframing the CUDA Moat: The Scale of Co-Optimization
To understand why alternative hardware struggles to gain traction, we have to look past the conventional narrative surrounding NVIDIA’s software prowess. CUDA’s real advantage isn’t simply software. It’s decades of continuous hardware-software co-optimization. There is a reason why more than half of Nvidia’s engineering workforce are focused on software development vs. hardware engineering alone.
The rest of the hardware market faces an existential challenge in trying to keep up with the frontier of advancements in AI architecture along with the maturity of the CUDA ecosystem. And there are likely roughly a thousand or so engineers globally that have the depth of expertise to do the joint software and hardware ML performance optimization required to extract maximum performance from new designs. To truly enable a heterogeneous hardware landscape, we must democratize co-optimization for other silicon players. Every chip architecture needs a scalable, high-velocity way to match the right kernels to the right workflows. Without it, alternative silicon remains functionally sidelined, regardless of its theoretical FLOPS.
This is where Infinity enters the stack. Infinity is building the automated software layer that makes any AI chip inference-ready, enabling hardware companies to compress years of systems engineering into days.
As more hardware platforms become practical for production AI, developers gain more choice, AI infrastructure becomes more resilient, and the cost of deploying AI can continue to decline.
Infinity’s core innovation is an autonomous optimization engine that searches the massive combinatorial space of potential kernel configurations to find the optimal fit for a specific model workflow on a specific target architecture.
By dramatically accelerating this discovery process, Infinity democratizes high-performance systems engineering. Early commercial engagements with partners like d-Matrix and others demonstrate this potential.
Leadership Built for the ML Systems Era
Solving one of AI infrastructure’s deepest challenges demands a rare synthesis of machine learning expertise and low-level systems engineering.
Jeremy Nixon embodies this rare intersection. Before founding Infinity, Jeremy was a researcher at Google Brain focused on AI-driven optimization. He also co-founded AGI House, one of the most influential communities in frontier AI research. Through both experiences, Jeremy developed a unique perspective on where AI infrastructure was headed and where the software layers were falling short.
Alongside Jeremy, the Infinity team brings together world-class expertise spanning compiler design, distributed systems, and hardware-level optimization.
At Touring Capital, we believe the future of compute belongs to both breakthrough hardware and the intelligent software layer that activates it. By automating one of the industry’s most complex engineering disciplines, Infinity is reshaping how AI hardware is brought to market. We are incredibly excited to partner with Jeremy and the entire Infinity team as they lay the software foundation for the next era of AI infrastructure.