Back to news
AI Tools & Products
21h ago

Goodfire introduces cost-effective internal monitors for AI safety

Oct 8, 2026
AI Summary

Goodfire has launched a new monitoring system that tracks internal processes of AI models, offering a more affordable alternative to traditional oversight methods. This innovation aims to enhance safety by detecting potential risks in real-time, particularly for open AI models that lack built-in safeguards.

  • Goodfire has developed internal monitors that observe AI models' internal workings rather than just their outputs, reducing costs associated with traditional AI oversight methods.
  • The monitors are available to customers of Baseten, which provides AI model hosting services, and are part of a safety partnership with Hugging Face.
  • The new system uses small detectors called probes to read internal signals during an AI agent's operation, flagging potential issues for further examination by a separate AI model.
  • Customers can customize monitoring for various risks, including hacking and misuse of biological weapons, and choose automated responses like logging events or human review.
  • Goodfire claims its monitoring approach is significantly cheaper, with costs for monitoring sessions being much lower than traditional methods, while also maintaining high detection rates for malicious activities.
  • The technology is particularly aimed at open models, which can be modified by developers to remove safety features, highlighting the need for effective monitoring in these environments.
  • Goodfire's research indicates that many leading open models exhibit vulnerabilities, with high rates of reward hacking detected in tests.
  • The company aims to further develop its technology to trace AI behavior back to its training origins, enhancing understanding and control over AI systems.
ai monitoringcost-effectiverogue agentsmodel inspectiongoodfire