The demo always works. You connect a model to the warehouse, ask “what was revenue last quarter?”, and watch it write SQL, run it and hand back a number with a tidy explanation. It looks like the analytics team just became optional.
In June, Anthropic’s own data team published how they actually run self-service analytics on Claude, and the first warning is about exactly that demo: “pointing Claude at a warehouse and letting the agents execute can create a false sense of precision.” After eight years in analytics, that sentence is the one I’d put on every AI-for-data pitch deck.
Data is not software
The post makes a distinction I think most enterprise AI plans miss. Coding has many acceptable answers, and code has guardrails: it compiles, tests pass or fail. Analytics usually has one correct answer, from one correct source. A query that runs and returns a plausible number is not a success. It’s the most dangerous kind of failure, because nobody checks it.
Anyone who has sat in a review where two teams brought two different revenue numbers knows this already. The hard part of analytics was never writing SQL. It was knowing which table, which definition, and which exclusions the business agreed on.
Three ways the agent gets it wrong
Anthropic names three failure modes, and they match what goes wrong with human analysts too, only faster.
- Ambiguity. “With hundreds of viable options in a data model (out of potentially millions of fields), the agent is unable to choose the correct fields that best answer a user’s question.”
- Staleness. Definitions and schemas change constantly, so the agent’s knowledge goes stale “and start[s] returning subtly wrong answers.”
- Retrieval. The right answer is documented, but “given the vastness of the search space, the agent simply doesn’t find it.”
None of these is fixed by a bigger model. They are all about what the model is allowed to know, and how reliably it finds it.
What the numbers say
The most useful thing in the post is that it measures. Without skills (folders of markdown that tell the agent how the team works), Claude “didn’t exceed 21%” on their analytics evals. With skills, sources of truth and validation, they report about 95% of business analytics queries automated at roughly 95% accuracy.
| Measure | Without skills | With skills |
|---|---|---|
| Accuracy | 21% | 95% |
Source: Anthropic, How Anthropic enables self-service data analytics with Claude (June 2026). 'With skills' is reported as consistently above 95%.
Then the part nobody puts in a demo. When they stopped treating the skills as living documents, accuracy drifted “from ~95% at launch to ~65% over a month.”
| Measure | At launch | One month later |
|---|---|---|
| Accuracy | ~95% | ~65% |
Source: Anthropic, same post.
Their fix was organisational, not technical: roughly 90% of their data-model changes now ship with a skill change in the same pull request. The documentation moves with the data, or it rots.
What worked, and what didn’t
The stack they describe has four layers:
- Data foundations: modelled tables, tests that run early, freshness and completeness checks.
- Sources of truth: a semantic layer, so that “if a question maps cleanly to a defined metric, the agent calls a function and gets one number, the same number every other surface in the company produces.”
- Skills: a thin router skill that points the agent at the semantic layer first and at about 30 reference files per domain second, and a runbook skill that works like a senior analyst: clarify the question, find the source, run the query, then have sub-agents argue with the result.
- Validation: offline evals, a before-and-after run for every meaningful skill change, and monitoring in production.
The failures are just as instructive. Letting an LLM generate metric definitions from raw tables and query logs didn’t work. Giving the agent raw search over thousands of past queries “moved accuracy by less than a point.” Rewriting the docs past a certain point made them “longer, not better.” And the adversarial reviewer that argues with every answer bought about 6% more accuracy for 32% more tokens and 72% more latency: worth it for a board number, maybe not for a quick question in chat.
“We recommend generating the documentation with Claude, but having a human own the definition.” That line is the whole essay.
The enterprise version of this problem
If this is what it takes inside the company that makes the model, with a data team this invested in evals, it’s worth being honest about what “connect Claude to our data” means elsewhere. The same pattern shows up in Anthropic’s engineering writing on tools: with many tools connected, “the most common failures are wrong tool selection and incorrect parameters,” and intermediate results pile up in the model’s working memory “regardless of relevance.”
So the question for an enterprise isn’t which model to plug in. It’s whether you have canonical datasets, definitions a human owns, and a way to measure when the answers are wrong. Anthropic’s own advice for starting is modest: “a handful of canonical datasets, a few dozen offline evals, and a thin knowledge skill will capture most of the upside.”
That’s not a model project. It’s an analytics project that happens to have a model at the end of it.
References
- How Anthropic enables self-service data analytics with Claude Anthropic (Claude blog), 3 June 2026.
- Introducing advanced tool use on the Claude Developer Platform Anthropic Engineering, 24 November 2025.