Goodfire deploys internal-signal probe monitors to catch risky AI agent behavior
Overview
Goodfire Research has built and deployed probe-based cyber monitors for Kimi K3 and GLM 5.3 that read a model's internal signals rather than only its written output.
A lightweight probe screens suspicious exchanges before an LLM judge reviews them. Goodfire says this cascade reaches about 93% recall with a 5.5% benign-session interruption rate, at roughly 50x lower judge cost. In its tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, catching 94% of malicious hacking sessions. These figures are Goodfire's own claims and have not been independently verified.
The company cites red-teaming by FAR.AI, its one independent element, which Goodfire says reduced universal jailbreaks to zero across 140 strategies and cut total jailbroken interactions by 97%. Goodfire describes these as results from a static battery of attacks.
TechCrunch reports the monitors are now available to customers of Baseten, a company that hosts and runs AI models for others. Customers can choose which risks to watch, including offensive hacking, chemical and biological weapons misuse, and reward hacking.
Written by AI from the articles below · updated Oct 8, 9:04 PM ET
Check the sources:
Developments
4 developments
- Oct 8, 12:29 PM ET · 1 articleGoodfire reports its monitor cut universal jailbreaks to zero in Far AI red-team testingGoodfire: .@farairesearch red-teamed our monitor with a static battery of attacks.
- Oct 8, 12:29 PM ET · 1 articleGoodfire releases cybersecurity monitors for Kimi K3 and GLM 5.3Goodfire: We built cybersecurity monitors for Kimi K3 and GLM 5.3 — more accurate, 50x faster, & 50x cheaper than an optimized LLM judge.
- Oct 8, 12:14 PM ET · 2 articlesGoodfire deploys probe-based cyber monitors for Kimi K3 and GLM 5.3 on a production inference stackGoodfire Research: Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade
- Oct 8, 12:00 PM ET · 1 articleGoodfire launches inside-out activation monitors for AI agents via BasetenTechCrunch · AI: Goodfire launches inside-out monitors to catch rogue AI agents at lower cost
Article timeline
Follow the coverage from different perspectives. Times are ET.
- Goodfire.@farairesearch red-teamed our monitor with a static battery of attacks.
AIIn this testing, our monitors cut successful universal jailbreaks to 0, and reduced total jailbroken interactions by 97%.
- GoodfireWe built cybersecurity monitors for Kimi K3 and GLM 5.3 — more accurate, 50x faster, & 50x cheaper than an optimized LLM judge.
AIExternal red-teaming from @farairesearch found that our monitor greatly reduces universal jailbreaks. 🧵
- Goodfire ResearchGoodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade
AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.
- TechCrunch · AIGoodfire launches inside-out monitors to catch rogue AI agents at lower cost
AIGoodfire has launched monitors that read a model's internal signals during agent work instead of reviewing its written output. The monitors are available to Baseten customers, who can choose risks to watch and set automated responses. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, while catching 94% of malicious hacking sessions.
Heat trend
Not enough continuous observations to show a trend yet.