Anthropic announced on September 18, 2026 that it is partnering with Accenture on embedded evaluation, an arrangement in which independent evaluators work inside Anthropic “with access comparable to an employee’s.” According to the post, “Anthropic and Accenture each expect to invest at least $1 billion in building capacity in this area over the next five years.”

On this page

What embedded evaluation means

Anthropic contrasts the model with how outside testing works today: “Unlike today’s external evaluators, embedded evaluators will work inside AI companies, with access comparable to an employee’s.” That access, the post says, lets evaluators observe model development during training, follow deployment decisions and work directly with staff.

The stated purpose is broader than testing individual models. “From this vantage point, embedded evaluators can assess how a company operates, verify that it is keeping its safety commitments, and identify blind spots,” Anthropic writes.

Who does the work

The partnership “will be led by Faculty, Accenture’s specialist AI business,” and will cover:

  • evaluating and red-teaming models;
  • conducting alignment assessments;
  • testing model safeguards.
Item Detail, per Anthropic
Partner Accenture, led by its AI business Faculty
Investment At least $1 billion each from Anthropic and Accenture over five years
Evaluator access “Comparable to an employee’s”
Scope Red-teaming, alignment assessments, safeguard testing, company operations
Other evaluators In dialogue with METR and other nonprofits on pilots

Anthropic adds that it is “in dialogue with METR and other nonprofit evaluators to pilot elements of embedded evaluation using their own funding.”

Where it comes from

The post presents the partnership as “an important step toward the commitment, made in our CEO’s essay ‘We Must Pace the Frontier,’ to embed evaluators within Anthropic.”

What it changes

By Anthropic’s own contrast with “today’s external evaluators,” the change is one of position: evaluators sit inside the company while models are trained and deployment decisions are made, rather than receiving access from outside. The arrangement also attaches a stated funding scale on both sides, gives Accenture’s Faculty the lead role, and leaves room for nonprofit evaluators such as METR to pilot parts of the approach with their own funding.

For related Anthropic safety measures, see our coverage of Anthropic’s alignment and security changes; Claude releases are in the AI model release timeline.

Sources