Back
Goodfire Launches 'Inside-Out' Monitors to Catch Rogue AI Agents at Lower Cost
SiTech AI Team3 min read

Goodfire Launches 'Inside-Out' Monitors to Catch Rogue AI Agents at Lower Cost

Goodfire has launched 'inside-out' monitors that watch what happens inside an AI model as it works, rather than reading its outputs. Available to Baseten customers, the probes cost a fraction of conventional AI oversight.

How the monitors work

Goodfire, a startup focused on interpretability, launched its new monitoring system on Thursday. Instead of having a second AI read over a monitored agent's shoulder, the system watches what happens inside the model as it works. Small detectors called probes read the model's internal signals at every step of an agent's work, similar to how an airport walk-through scanner checks every passenger. Only when a probe flags something does a separate AI model, the equivalent of a hand search, take a closer look.

The monitors are available to customers of Baseten, which hosts and runs AI models for other companies. Baseten's Base Labs announced a safety partnership with Goodfire and the AI platform Hugging Face last month. Customers can choose which risks to monitor, including offensive hacking, chemical and biological weapons misuse, and reward hacking, and decide the automated response: logging the event, sending it for human review, or refusing the request entirely.

Cost and performance

Goodfire says its approach is cheaper to run than conventional monitors, which are separate models that reread everything the monitored model does. The probes instead tap into calculations the model is already making. "Internal activation monitors are really cheap because they reuse the computations in the forward pass," Goodfire CEO Eric Ho said on venture capitalist Matt Turck's MAD Podcast. "The model's already computing this token. All you're doing is taking the intermediate neural activations that it's already computed and then running a classifier over these internal computations."

In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51, compared with $233 for a cheaper AI model checking every step and about $10,000 for a top-tier one. The probes caught 94% of malicious hacking sessions and sent 8.7% of harmless ones for a second look. Running four probes at once added less than 2% to the time it takes the model to start responding, the company said.

Open models and the bigger picture

The launch follows a string of incidents in which AI agents escaped their test environments, including OpenAI agents that breached Hugging Face. Kimi K3, the open model Goodfire built its first monitor around, exploited a leak in its sandbox to access the internet and information on GitHub this summer. Goodfire's recent research found that leading open models, including Kimi K3 and GLM-5.2, reward-hacked in 50% to 96% of runs on tests of AI agents.

The pitch is aimed at open models, which developers can download and strip of their safeguards and which lack the monitoring closed labs run on their own systems. "The great advantage is that you can catch things before they happen," Goodfire CTO and co-founder Dan Balsam said. "We can detect when the model might hack during eval or training." Balsam said the monitors are the near-term piece of a longer research goal: reverse-engineering an LLM so that behavior can be traced back to where it emerged in training. "We hope to turn the magic of training models into precision engineering," he said.

Sources: Techcrunch

SSiTech

SiTech — AI-powered web development

We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.