//
Discussion · 2026-09-13

Embedded AI evaluators: real oversight or incumbent liability theater?

Dario Amodei’s weekend “pace the frontier” push reduced to its one tangible outcome: embedded third-party…

The Debate

For vs against
Essence
Dario Amodei’s weekend “pace the frontier” push reduced to its one tangible outcome: embedded third-party evaluators inside the frontier labs. Gavin Baker argues the regime is smart liability management — there is no Section 230 shield for model outputs — with minimal investment implications and a “smoother for longer” cycle. Skeptics counter that lab-chosen, lab-housed evaluators are independence theater that entrenches incumbents, just as a White House executive order takes shape.
For
  • Embedded third-party evaluators are smart liability management: there is no Section 230–style shield for model outputs, and demonstrating a “duty of care” will matter in future litigation — several internet companies might have gone bankrupt without Section 230. (Gavin Baker)
  • The single tangible new fact has minimal investment implications; for anyone wanting a “smoother for longer” cycle, most constraints are good — wafers, watts, real rates and spreads — and we are nowhere near excessive regulation. (Gavin Baker)
  • The industry is converging on oversight models anyway: Elon’s MPAA-like competitor peer review, Demis’s FINRA-like self-regulatory structure, Sam Altman’s agreement to implement embedded evaluators — the evaluator regime is the compromise form. (Gavin Baker, summarizing)
  • The read has been “balanced, thought through and well articulated, without the histrionics” of the loudest voices — supporters treat Baker’s framing as the reasonable center of an emotional weekend. (Ariel Szin; “Good summary” — Sriram Krishnan)
Against
  • “Embedded evaluators chosen by the lab, sitting at desks the lab provides. That’s not independence. That’s a liability exhibit” — the regime directly inverts Baker’s duty-of-care thesis. (Jess)
  • The idea is “fraught with a host of associated problems” in the real world; global cooperation is unworkable, and “Washington’s number one goal will be to weaponize it for political gain.” (Bill Carson)
  • “Too much power in the hands of too few companies with oversight provided by third parties who may be biased” — evaluators must not be affiliated with any lab. (Martha Johnson; Maureen O’Toole)
  • Dario saying he’d hand Anthropic to “the right mix of governments” reveals the end-state is state entanglement with the labs, not safety — and it reframes everything the regime touches. (Poor Richard)
  • The incentives cut the other way: “The incentives are not aligned for the benefit of most of humanity,” and if labs are truly far ahead internally, the pacing push is an admission that capability — not safety process — is what’s being managed. (wayne nelms; BrianL)

View the thread on X

The Tweets

4 posts
GBGavin Baker@GavinSBaker · 4:17 PM · Sep 13, 2026X Post

Wild 24 hours for AI and lots of different proposals have been made. TLDR; the only *tangible* new fact is that OpenAI and Anthropic are going to have embedded 3rd party evaluators from unknown organizations with Dario floating METR as a possibility. Having 3rd party evaluators is smart as there is no Section 230 style liability shield for model outputs and showing a "duty of care" will be important in future litigation. Several internet companies might have gone bankrupt without Section 230 so limiting liability really matters. There are minimal investment implications from this single new fact, but I do think that for anyone who wants a "smoother for longer" cycle then most constraints are good: wafers, watts, real rates and spreads. Excessive regulation is a different matter but I don't think we are anywhere close to this even if the vector changed over the last 24 hours. To summarize the events: Dario made the most maximalist proposal of the weekend: embedded 3rd party evaluators, a national regulatory regime for models beyond a certain capability/ingredient threshold, a broad international regulatory pact between democracies, stricter limits on compute/distillation for China and then a different international regulatory regime that encompasses China. Before there is a national regulatory regime, he wants a Sherman act waiver so that Anthropic can safely coordinate with OpenAI and other frontier labs without antitrust fears. TBF, this latest proposal is much less maximalist than some of his prior proposals like "Policy on the AI Exponential," where he advocated for an FAA for AI. I believe he is sincere in his beliefs. And despite all the protestations, all of this would also probably be good for his business over the long-term. Sam agreed that embedded 3rd party evaluators were a good idea and stated they would implement them. Again, this is smart as should help limit future liability. Elon said "Dario is right" and later specified that "Dario is right that there should be some oversight. Peer review of AI by competitors is the right way to start this off." This would be a MPAA like self-regulatory structure for AI with regular calls between the labs plus a process where each new model is evaluated for safety by competitors for a 1-2 week period before being released. That is *wildly* different from Dario's proposal and in-line with what David Sacks has been proposing. Elon also stated that nothing was going to slow down open-weight models. Demis said that Dario's essay was a "step in the right direction." Dario also said that he was also open to Demis' idea of a FINRA like self-regulatory structure as part of his proposal. David Sacks had a thoughtful post where he said that Dario and Sam should pace unilaterally, called the antitrust waiver a cartel request and denied that METR was truly independent given their ties to Anthropic. Sriram Krishnan, former White House AI advisor, noted that it would be important to have the 3rd party evaluators come from independent organizations that are not affiliated with any lab, which is basically an indirect statement about the relationship between METR and Anthropic which Sacks was explicit about. Clem from Hugging Face said they were open to being a neutral 3rd party evaluator, which is interesting especially if Jensen was consulted before that post. Alexander Wang from Meta noted that alignment would be an increasing focus going forward. An executive order seems likely after all this and the language in this EO is going to be really important. It is possible to democratize and distribute AI broadly and safely without centralizing it in the hands of a few corporations who might each become more powerful than any single government. I do not want a few humans in control of intelligence. I want us all to have our own intelligences that reflect our own values and human variation in all of its richness. Intelligence distribution over intelligence centralization FTW.

524 Reposts · 2865 Likes · View on X →
JEJess@Jesscarson · 2hX Post

Embedded evaluators chosen by the lab, sitting at desks the lab provides. That's not independence. That's a liability exhibit

1 Reposts · 7 Likes · View on X →
BCbill carson@BillCarson3x · 6:14 PM · Sep 13, 2026X Post

third-party embedded evaluators may sound great on paper in the real world this idea is fraught with a host of associated problems. As for global cooperation? you can throw that out the window and there's no way to dance around that, there's no regulatory process that can move the needle here. Then your next biggest problem is Washington's involvement. In the new state of politics we find ourselves in Washington's number one goal will be to weaponized it for political gain and if you think otherwise you're smoking crack out of a Pepsi can

PRPoor Richard@Bnjmn_frnkln76 · 15 min agoX Post

Pitch perfect as always. Curious how your thought process changes with Dario saying he'd be open to handing Anthropic over to, "the right mix of governments."