Deep Cogito Raises $43 Million Series A for AI Post-Training
  • News
  • North America

Deep Cogito Raises $43 Million Series A for AI Post-Training

Funding will support reinforcement learning research, infrastructure, and enterprise AI expansion.

8/27/2026
Ghita Khalfaoui
Back to News

Deep Cogito, a San Francisco-based artificial intelligence research company focused on post-training and reinforcement learning, has raised $43 million in Series A funding to expand its work on systems designed to improve the reasoning capabilities of advanced AI models. The round was led by TQ Ventures, with participation from Benchmark, Nexus Venture Partners, Atreides Management, South Park Commons, and cybersecurity company Zscaler. The financing brings Deep Cogito’s total funding to more than $56 million and will support research, engineering, infrastructure, and enterprise expansion.


Building on Google AI Search Experience

Deep Cogito was founded by Drishan Arora and Dhruv Malrana, who previously worked on Google’s AI-powered search products, including AI Mode and AI Overviews. During their time at Google, Arora led post-training work involving Gemini models for AI Search, while Malrana was responsible for developing the product from its early stages. Their new company is based on the view that post-training will become increasingly important as developers seek to turn pre-trained foundation models into stronger reasoning systems capable of handling more complex tasks.

Advancing AI Through Post-Training

The company’s research centers on large-scale reinforcement learning and techniques intended to enable AI systems to progressively improve their own capabilities. One area being explored is Iterated Distillation and Amplification, or IDA, an approach in which models are given additional computational resources to generate stronger answers before those improvements are incorporated back into their underlying parameters. Deep Cogito ultimately aims to develop systems that can continue advancing beyond the limitations imposed by relying primarily on human-generated training data.

Developing the Cogito Model Family

Deep Cogito has tested its post-training techniques through its Cogito family of open-weight models, which have been released across sizes ranging from approximately 3 billion to more than 600 billion parameters. According to the company, these releases demonstrated its ability to enhance existing models using reinforcement learning and other post-training methods at a scale typically associated with major artificial intelligence laboratories. The same underlying technology is now being adapted for enterprises seeking specialized AI models trained around proprietary information, business decisions, and measurable outcomes.

Enterprise Applications Expand

The company is positioning its technology as an alternative to lighter forms of model customization for organizations that require more specialized artificial intelligence capabilities. Zscaler, which initially began working with Deep Cogito as a customer, participated in the Series A as a strategic investor after using the company’s technology to develop models aligned with its security products and performance requirements. Deep Cogito said its approach allows businesses to incorporate domain-specific knowledge and objectives directly into model training rather than relying solely on prompts or relatively limited customization layers.

Funding Supports Research and Infrastructure

Deep Cogito plans to use the Series A capital to expand its research and engineering teams while increasing the computing infrastructure required to train and improve frontier-scale models. Funding will also support future releases in the Cogito model family and the expansion of Deep Cogito’s work with enterprises seeking to build specialized intelligence using their own proprietary datasets. Investors backing the round see post-training as an increasingly significant layer of the AI technology stack as companies seek greater control over model performance, specialization, and continuous improvement.


The $43 million Series A gives Deep Cogito additional resources to pursue its strategy of developing reinforcement learning and self-improvement systems for increasingly capable artificial intelligence models. By combining open-weight model research with customized enterprise deployments, the company is attempting to demonstrate that advanced post-training techniques can be applied both to frontier research and commercial use cases. Its next phase will focus on scaling those capabilities while testing whether self-improving training approaches can deliver sustained advances beyond what conventional model development methods currently provide.