Nebius Acquires Inferize to Boost AI Inference Efficiency
  • News
  • Europe

Nebius Acquires Inferize to Boost AI Inference Efficiency

Inferize joins Nebius Token Factory to improve GPU utilization and model scaling

10/2/2026
•Ghita Khalfaoui
Back to News

Nebius, the AI cloud company listed on Nasdaq under the ticker NBIS, announced on October 1, 2026, that it has acquired Inferize, an inference optimization company. The acquisition brings technology that shortens the time required to launch and scale large AI models. Inferize's team and assets will join Nebius Token Factory, the company's managed inference platform for production AI.


Closing the Inference Readiness Gap

In production AI, cold starts occur when models need time to load before serving a single request. This delay leaves assigned GPUs idle at launch time, during demand spikes, and when model weights are updated mid-run, including during reinforcement learning. Platforms often hold spare capacity to meet service-level targets, which increases costs and worsens token economics.

Integration with Nebius Token Factory

Inferize's technology addresses what the company calls the idle GPU tax by enabling capacity to scale more closely with actual usage. This supports higher utilization rates and better token economics for production AI workloads. Nebius plans to integrate Inferize's technology directly into Token Factory, making the platform more responsive to customer demand.

Executive Perspective

Danila Shtan, Chief Technology Officer of Nebius, said that running inference well requires more than fast GPUs and optimized models. He explained that the entire system must respond when demand changes, especially how quickly additional capacity becomes ready to serve customers. Shtan added that Inferize brings both accelerating technology and deep expertise in GPU systems into Token Factory.

Team Background and Mission

Guy Bortnikov, co-founder and CEO of Inferize, said keeping spare GPUs running is the price of being ready for demand. Removing that cost is what the company was built to do, and Nebius is where the technology can go directly into the platform. The Inferize team will work across Nebius Token Factory with the objective of serving more customer demand from every GPU.

Strategic Context

Inferize was founded in January 2026 and had a working prototype within three months. The acquisition adds another layer to Nebius's production inference strategy, following Eigen AI's optimization work at the model, kernel, and system levels. It also complements Clarifai's core team and licensed technology for system-level inference and compute orchestration.

Broader Industry Relevance

The acquisition highlights how inference optimization has become a central focus for AI cloud providers as model deployment matures. Enterprises running large AI models increasingly face cost pressures related to idle compute and startup delays, and Nebius is positioning Token Factory to compete by addressing these operational inefficiencies directly. This approach moves beyond raw hardware performance to improve the total economics of serving AI workloads.

Transaction Terms

The terms of the transaction were not disclosed, and Nebius did not provide additional financial details about the acquisition. The deal adds deeply experienced systems engineering talent to Nebius's inference team. It also continues the company's production inference build-out with optimizations across the stack, supporting its broader AI cloud strategy.


The acquisition of Inferize strengthens Nebius Token Factory's ability to handle production AI workloads more efficiently. By reducing idle GPU time and improving token economics, the integration supports Nebius's goal of maximizing useful work from its infrastructure. As demand for large-scale AI inference continues to grow, the combined team is positioned to deliver more responsive and cost-effective AI cloud services.