Base Labs Launches Open-Weight AI Safety Partnership with Hugging Face and Goodfire
  • News
  • North America

Base Labs Launches Open-Weight AI Safety Partnership with Hugging Face and Goodfire

The initiative targets abliterated models with open-weight evaluation and monitoring infrastructure.

9/17/2026
Ali Abounasr El Alaoui
Back to News

Base Labs, the research arm of AI infrastructure provider Baseten, announced a new open-weight AI safety partnership with Hugging Face and Goodfire AI on Wednesday. The collaboration aims to create evaluation and monitoring infrastructure that supports safer training and deployment of open models. The move comes as industry debate intensifies over how to preserve openness while preventing models from becoming harmful.


Addressing Open Model Risks

The initiative directly addresses the rising technique known as abliteration, which removes built-in safeguards from open-weight models. Hugging Face, a major repository for open source AI models, currently lists more than 6,000 abliterated models. This scale has pushed model safety from a research concern into an urgent operational challenge for developers and enterprises.

Safety Across the Model Lifecycle

Base Labs will develop and publish methods for training, evaluating, and monitoring open models. Baseten plans to integrate these research outputs into its deployment infrastructure, enabling live runtime monitoring for models served on its platform. The partners frame the effort as a transparent standard embedded from the start rather than a layer added after release.

Complementary Roles Across the AI Stack

The three organizations bring complementary strengths to the collaboration. Hugging Face serves as the distribution hub where open models are shared and downloaded, while Goodfire contributes expertise in interpretability and explaining model decision making. Baseten acts as the serving and infrastructure layer, connecting safety research to production deployment for customers.

Substantial Financial Backing

The partners enter the effort with considerable financial resources. Baseten raised a $1.5 billion Series F in June, pushing its valuation to $13 billion. Goodfire secured a $150 million Series B earlier this year, reflecting investor confidence in its model interpretability platform and the broader safety mission.

Why Interpretability Matters

Goodfire's interpretability work is central because understanding model internals can help identify unsafe behavior before it spreads. If researchers can detect the internal patterns associated with harmful outputs, they can design better monitors and interventions. This technical visibility is especially important in open ecosystems where user-facing filters alone are not sufficient.

The Challenge of Downstream Modification

A key difficulty is that open-weight models can be altered after release, bypassing safeguards that worked in the original version. Once a model is downloaded and modified, the organization that released it has limited control over how it is used. This creates a distributed responsibility problem that the partnership seeks to address through monitoring and shared standards.

Growing Industry Alignment

The partnership follows broader industry moves toward more coordinated open model safeguards. Earlier in 2026, Nvidia led the formation of the Open Secure AI Alliance, indicating that hardware and infrastructure players see safety as foundational. Baseten's new effort fits into this pattern by connecting safety research directly to commercial deployment rails.

An Open Call to the Ecosystem

Baseten is inviting outside developers to contribute to the framework, signaling that it wants the initiative to become ecosystem wide rather than owned by one vendor. If it gains traction, the work could influence how model builders and serving platforms approach safety in open-weight AI. The announcement highlights a broader shift toward treating safety as a shared technical layer.


The partnership signals an important evolution in how the AI industry handles open-weight model security. By combining model hosting, interpretability, and inference infrastructure, the group hopes to make safety measurable, transparent, and practical across the full model lifecycle. Instead of treating openness and safety as opposing forces, the collaboration positions them as mutually reinforcing priorities for the next phase of AI development.