动察 Beating AI News: Laminar, an AI Agent observability platform, has released flow-1, a model specifically designed to inspect Agent execution traces. It reads model calls, tool calls, and returned results to identify where and why an Agent made a mistake.
flow-1 runs within Laminar's Signals agent. The model can search an entire trace (the complete execution record of an Agent) and then inspect specific steps as needed. During training, Laminar first used synthetic investigation data for supervised fine-tuning, then trained tool calling and complex trace analysis through reinforcement learning.
In Laminar's self-built test set of 523 difficult traces, flow-1 achieved an error detection F1 of 0.835, while GPT-6-sol scored 0.816. sol has higher recall and can find more real errors; flow-1 has higher precision and produces fewer false positives.
For traces under 100,000 LLM tokens, flow-1 costs about $0.0011 per analysis on average, compared with about $0.026 for GPT-6-sol. According to Laminar's estimates, $1 can analyze about 888 and 38 traces respectively, a difference of about 23 times. flow-1 is also about 25% cheaper than GPT-6-luna.
At present, these results all come from Laminar's self-built benchmark, with no third-party retesting yet. flow-1's training data mainly comes from synthetic workflows, about 48% of which are software engineering tasks. Some in the community have already questioned whether it can maintain the same performance when facing novel failures in real production environments that were not pre-designed.

