Beating AI News Flash: Anthropic has officially begun implementing "embedded evaluation." Accenture's AI company Faculty will send personnel into Anthropic to evaluate and red-team test models, conduct alignment assessments, and inspect safety measures.
This is very different from ordinary external evaluation. Evaluators will receive near-employee-level access, be able to view the model training process, understand development and deployment decisions, and communicate directly with internal employees.
Over the next five years, both sides will each invest at least $1 billion to build this evaluation capability. Anthropic also acknowledges that the industry currently lacks even unified rules, and what exactly evaluators can see, how results will be disclosed, and who will pay are still being determined. For now, Accenture's evaluation work is paid for directly by Anthropic.
But Accenture itself is a long-term strategic partner of Anthropic, and has even established a dedicated Anthropic Business Group to train about 30,000 people to use Claude. One of the minimum conditions proposed by the AI Evaluator Forum on the same day is precisely that evaluation institutions should not have significant commercial dealings with the companies being evaluated.
Anthropic will later bring in other evaluation institutions, and is also in talks with nonprofit organizations such as METR (also a longtime partner of Anthropic) about pilots.

