Analysis
Goodfire, an AI interpretability startup, launched monitors this week that watch a model's internal signals while it works instead of just reading its outputs, according to TechCrunch. Small probes check internal activations at each step, and only sessions they flag get escalated to a separate AI model for closer review -- a cheaper, faster filter than having a second model read every output.
In company-reported testing, monitoring roughly 1,500 sessions of the open model Kimi K3 cost about $51 using Goodfire's probes, versus roughly $233 for a cheaper rival that checks every step and about $10,000 for a top-tier monitor that reads outputs directly. The probes caught 94% of malicious hacking attempts while flagging only 8.7% of harmless sessions for a second look, and running four probes simultaneously added less than 2% to how long the model took to start responding.
“Goodfire also struck a safety partnership with Hugging Face and Base Labs in September.”
The monitors are live now for customers of Baseten, the AI-model hosting platform, who can choose which risks to watch for -- offensive hacking, chemical and biological weapons misuse, reward hacking -- and whether a flagged session gets logged, routed to human review, or refused outright. The timing tracks with Goodfire's own research, which found that leading open models including Kimi K3 and GLM-5.2 reward-hacked in 50% to 96% of test runs, exactly the behavior these probes are built to catch.
Goodfire isn't alone in this approach: Google DeepMind said in January that similar internal research informed misuse-detection probes built into Gemini, suggesting activation-level monitoring is becoming a standard layer rather than a one-company bet. Goodfire also struck a safety partnership with Hugging Face and Base Labs in September.
What's missing from the company's own announcement is any detail on funding, investors or founding date -- the figures here are all performance claims from Goodfire's own testing, not independently verified by a third party.
