The work moved to AI. The platforms did not.
Nothing in a data platform was designed to be handed to something that makes hundreds of changes unwatched.

Everything changed except the bottleneck
Agents are already writing data pipelines. Within three months of launch, 60% of new pipelines on Databricks’ LakeFlow were agent-written (Databricks, DAIS 2026). Over the same period, real-world agent incidents grew 4.9× month on month (CLTR and the UK AI Safety Institute, 2026).
Meanwhile 7% of enterprises say their data is completely ready for AI to consume (Cloudera and HBR, 2026). That figure is about the data, not about whether a platform is ready to let an agent change it — which is the harder question, and the one nobody is measuring.
Not one of them was a model failure
The incidents that reached production read as a list of missing containment, not of bad reasoning:
- Replit wiped a production database during a declared code freeze, then generated records to cover it. No rehearsal. Fortune, 2025
- Railway found a token on disk and deleted a production database and its backups in nine seconds. Full access. Cerbos, 2026
- Meta shipped a Sev 1 where an agent published unverified advice and exposed company and user data. Nobody checked. The Verge, 2026
- Uber burned an entire year’s AI budget in four months on an uncapped rollout. No limit. Forbes, 2026
Anthropic reviewed 141,006 evaluation runs in which Claude could have reached the internet, and found three where it gained access to the real systems of other organisations. Its own description of what went wrong was “closer to a harness and operational failure than a model alignment failure.” Anthropic, 2026 The lab that builds the model reached the same conclusion in public, about itself.
So a better model is not the answer
If none of those was a reasoning failure, none of them gets fixed by a stronger model. Each one is a gap in the world the agent was working in: somewhere to be wrong that is not production, limits that hold when the agent reasons its way past an instruction, evidence that a change is correct, and a ceiling on what it can spend. A general coding agent brings none of those to data work, and a data change that breaks nothing loudly gives nobody a reason to look.
That is why we think the harness is the product. Models improve for everyone equally, which makes them a commodity input. What does not transfer is the domain expertise a team encodes around them: the curated skills, the chosen workflows, and the thousands of small editorial decisions that make a harness useful in context.
The shape of the answer
Vibedata is what that argument builds to: a data engineering agent you can trust with production data, because of what surrounds it rather than what powers it. Three things carry the weight, and the model does the rest.
- Knowledge. Every decision the agent makes lands in a file you own.
- Isolation. A zero-copy clone of production the agent can be wrong in.
- Skills and tools. The data judgment is written down rather than hoped for, and you can add your own.
Underneath them, the same three on every data platform, so the argument does not become a feature of one vendor. Your team stays the accountable driver: the work is delegated inward, and a human has no role inside the agent loop but every role at its edges.
Where this goes next
Applied to platforms rather than to teams, the same argument is what ADER sets out: what a data platform has to give an agent before it can run one at all.