AI interpretability company Goodfire has developed cybersecurity monitoring systems for Kimi K3 and GLM 5.3 that detect potentially harmful agent behavior while significantly reducing monitoring costs. Announced on October 8, 2026, the technology combines lightweight neural network probes with large language model judges to identify suspicious activity during inference. Goodfire reports approximately 93% detection recall at a 5.5% benign-session interruption rate, with monitoring costs up to 50 times lower than conventional approaches.
A New Approach to Cybersecurity Monitoring
Monitoring autonomous AI agents presents particular challenges because legitimate cybersecurity research and malicious exploitation often involve similar technical activities. Traditional monitoring methods rely on separate language models to examine interactions, creating substantial computational expenses and delays during extended agent sessions. Goodfire's approach instead examines internal model activations to identify potentially dangerous behavior before an agent executes external actions.
The monitoring architecture uses specialized probes as an initial detection layer, continuously evaluating model activity for signals associated with cybersecurity misuse. Interactions exceeding calibrated risk thresholds are forwarded to an LLM judge for further assessment, while ordinary activity proceeds without additional scrutiny. This selective evaluation reduces unnecessary processing while preserving the ability to intervene when potentially harmful behavior emerges.
Production Deployment and Infrastructure
Goodfire integrated its monitoring technology into SGLang, an inference framework supporting high-throughput language model deployment. The implementation incorporates custom computational kernels and compatibility with infrastructure features including speculative decoding, prefix caching, and distributed inference. Testing across four production-style workloads showed throughput differences below 2%, which the company considers within normal benchmark variation.
The system also substantially reduces monitoring delays compared with evaluating every interaction through a separate language model. Without judge reasoning, average additional latency declined from 160 milliseconds to approximately 2 milliseconds per agent turn. With reasoning enabled, the corresponding delay fell from 2,750 milliseconds to 42 milliseconds.
Detection Performance and Cost Efficiency
Goodfire evaluated its monitoring architecture using approximately 2,400 cybersecurity and general-purpose sessions encompassing more than 60,000 interaction turns. The assessment included simulated attacks, realistic agent environments, and real-world conversations, with Claude Opus 5 providing reference labels. At the selected operating threshold, the monitoring cascade detected approximately 93% of harmful sessions while interrupting 5.5% of benign sessions.
Cost comparisons demonstrated significant efficiency improvements over conventional LLM-based monitoring configurations. Processing one million agent turns cost approximately $65 without judge reasoning and $185 with reasoning, compared with $3,166 and $5,420 for their respective standalone judges. These results represent reductions of approximately 49 times and 29 times, respectively, in estimated judge costs.
Independent Security Evaluation
Independent AI safety organization FAR.AI conducted preliminary testing of Goodfire's deployed Kimi K3 monitoring system using a fixed collection of jailbreak attacks. Across 140 selected attack strategies, the monitors eliminated successful universal jailbreaks, although some individual harmful interactions still succeeded. The evaluation recorded 18 successful jailbroken interactions after monitoring, compared with 700 without the monitoring system.
Goodfire also tested the architecture on GLM 5.3 using training data originally developed for Kimi K3. In full-session evaluations, harmful-session blocking increased from approximately 29% to 96% for Kimi K3 and from 36% to 91% for GLM 5.3. However, harmless-session interruption rates also increased, highlighting the continuing trade-off between stronger safeguards and unrestricted legitimate activity.
Goodfire's monitoring architecture demonstrates a potentially practical approach to securing autonomous AI systems without imposing substantial inference costs or computational delays. Its results suggest that activation-based monitoring can complement traditional safety mechanisms while enabling interventions before potentially harmful tools are executed. Nevertheless, further adaptive security testing and broader deployment evaluations will be important to establish its effectiveness against evolving cybersecurity threats.