The work moved to AI. The platforms did not.

Nothing in a data platform was designed to be handed to something that makes hundreds of changes unwatched.

Many data sources narrowing through a single gate before reaching a governed pipeline into the warehouse
The agent scaled. The platform around it did not.

Everything changed except the bottleneck

Agents are already writing data pipelines. Within three months of launch, 60% of new pipelines on Databricks’ LakeFlow were agent-written (Databricks, DAIS 2026). Over the same period, real-world agent incidents grew 4.9× month on month (CLTR and the UK AI Safety Institute, 2026).

Meanwhile 7% of enterprises say their data is completely ready for AI to consume (Cloudera and HBR, 2026). That figure is about the data, not about whether a platform is ready to let an agent change it — which is the harder question, and the one nobody is measuring.

Not one of them was a model failure

The incidents that reached production read as a list of missing containment, not of bad reasoning:

Anthropic reviewed 141,006 evaluation runs in which Claude could have reached the internet, and found three where it gained access to the real systems of other organisations. Its own description of what went wrong was “closer to a harness and operational failure than a model alignment failure.” Anthropic, 2026 The lab that builds the model reached the same conclusion in public, about itself.

So a better model is not the answer

If none of those was a reasoning failure, none of them gets fixed by a stronger model. Each one is a gap in the world the agent was working in: somewhere to be wrong that is not production, limits that hold when the agent reasons its way past an instruction, evidence that a change is correct, and a ceiling on what it can spend. A general coding agent brings none of those to data work, and a data change that breaks nothing loudly gives nobody a reason to look.

That is why we think the harness is the product. Models improve for everyone equally, which makes them a commodity input. What does not transfer is the domain expertise a team encodes around them: the curated skills, the chosen workflows, and the thousands of small editorial decisions that make a harness useful in context.

The shape of the answer

Vibedata is what that argument builds to: a data engineering agent you can trust with production data, because of what surrounds it rather than what powers it. Three things carry the weight, and the model does the rest.

Underneath them, the same three on every data platform, so the argument does not become a feature of one vendor. Your team stays the accountable driver: the work is delegated inward, and a human has no role inside the agent loop but every role at its edges.

Where this goes next

Applied to platforms rather than to teams, the same argument is what ADER sets out: what a data platform has to give an agent before it can run one at all.

See how Vibedata works