Data Copilot is an agent product for doing data work, with a chatbot web UI as its entry point. It is becoming a data platform built around an AI data agent.

We started it at a hackathon in November 2025. The thought at the time was: since Codex/Claude Code are this capable, what if we designed the directory (the workspace) they work in, and built a product on top of that? The hackathon went well. By March 2026 I finally found a solid block of time, and over roughly two weeks to a month, together with another partner, I built the first version. Then we mentioned it in the company Slack. No mandate, no large-scale marketing.

We had already tried several times to build AI data products, none of them successful, so this time we did not assume it would see fast adoption — and we were unwilling to force people to use something for the sake of using it, so our expectations were not high. Until one day my collaborator said to me: "It seems like a lot of people are using our Data Copilot!"

I remember we checked the backend at the time: the previous day there had been a dozen or so users, asking a few hundred questions in total. That gave us some confidence. And then, in exactly those days, we discovered a bug in the product — a bug that made users wait one or two minutes for the first character of a response, with a full response taking something like five to ten minutes.

But looking at the backend, those dozen-odd users were still diligently interacting with it, a few hundred questions a day, averaging dozens per person.

That was the moment I knew this product was going to work. I said to my collaborator: we have finally seen what a successful product looks like, and a successful product is one where as long as the core value genuinely holds, users will put up with unimaginable inconvenience to use it. Because if it were me using this product, at that speed, I would have smashed my phone and sworn never to give it another chance.

Which raises the question: all of these users have Codex and Claude Code. Why are they still using this thing?

Users Really Are Using It Intensively and Voluntarily, When They Have Other Options

Last month:

Retention is extremely high, with no decay. That means users are using this product intensively and continuously.

We also ran a small survey, with 25 respondents: 60% said they would be "very disappointed" if Data Copilot disappeared, and 88% answered either "very disappointed" or "somewhat disappointed."

We did not force anyone to use this product, and we never ran heavy promotion. Internally it has also developed along free-market principles — that is, we have never tried to lock it in organizationally (as an organizational decision to use this and not something else).

I say this not out of a pursuit of some illusory "fairness," but because:

First, in an era of this much change, using an organizational decision to fix the "status" of an internal product does not mean much. If it is no good, it will be eliminated soon enough.

Second, only in a relatively free-market environment can we get real signal to guide the product's iteration. If we could force users to use it, we would only need to think about how to force them — and this does sometimes happen in enterprise software.

And of course I also hold a personal value: I want the things I build to make users' lives easier when they use them. That accumulates good karma; otherwise it accumulates resentment.

At the same time, I should account for the other possible explanations:

  1. More than half of our users have Codex/Claude Code, and many of them use both. Data Copilot also offers integration with these general-purpose agents (command line and MCP), so users are choosing our product from between general-purpose agents and Data Copilot.
  2. For cost-control reasons, we later required users to connect their own Codex, and we run their requests on their subscription. Under those conditions usage did not drop at all, so users are not using Data Copilot because of free tokens.
  3. Among users who use the command-line and MCP integrations, most still use the web UI as well; moreover, the questions they ask through the command line and MCP are not fundamentally different from those they ask through the web — both are dominated by analytical and engineering tasks of considerable complexity and business meaning, and rarely simple data pulls. So users are not using Data Copilot merely because it has integration capabilities other agents lack.

In sum: in competition with general-purpose agents, Data Copilot is the agent product users have chosen, and use intensively, as their entry point for data analysis and data work. We can rule out simple permission and access friction as the reason: it is not because they lack Codex/Claude Code, not because of free tokens, and not simply because Data Copilot has data-access capabilities that general-purpose agents lack and cannot integrate.

So why do users use Data Copilot?

Because Data Copilot Is the Product for Doing Data Work

When Manus was popular, I heard someone say: "Isn't this just a Claude Code wrapper?" The same view gets applied to Data Copilot. Manus and Data Copilot are indeed both wrappers around general-purpose agents, literally. But that is a technical perspective, an implementation perspective — not the user's perspective.

And what is the user's perspective? It is: under what circumstances do I think of this name; when I do, what picture appears in my head; what do I expect to get from it; and can that expectation be met.

For Manus users, who cares whether Manus is a Claude Code wrapper? Manus is the thing where: you open an elegant chat input box, you arrive at the office at 8:30 in the morning with your coffee and tell it, generate me a briefing report on $TSLA, here is our internal data spreadsheet, here is the previous template; then half an hour later you come back and it is 80–90% done, you give it a few more prompts and it fixes things, and that afternoon you walk into the meeting with that report. Could Claude Code do it? Yes, it could — maybe with 5x the effort.

Likewise, to its users Data Copilot is not some Codex wrapper; it is that input box on the light green web page. I wake up in the morning and ask it, send me that sales report, I want to look at the trend — and you get a Metabase report link, with a table and chart inlined in the chat below it; then I say, add me a filter, I want to see it by channel, and it gives you a channel breakdown.

So what is a product? A product is a physical or digital apparatus that carries a particular cognitive form and its corresponding functional expectations. Which tells you exactly what is not a product — capability is not a product. And what users want is a product, not a capability.

In the Data Copilot example above, the general-purpose agent has every one of these capabilities, but it is not that product.

First, the way it introduces itself to the user is not "I do data." Why would I go into something with "Code" in its name to do "data analysis"? A user's mind cannot naturally connect a general-purpose agent with doing data work, and that produces friction at the very first step.

Second, its context is not designed for data work. You ask it for a sales report and it may go find you the written report the sales department wrote — you may not like that, but that is what it is expected to do; by definition, it is supposed to do the thing most likely to satisfy your expectations within a broad workspace. You ask it how the sales trend looks, and it may tell you in a couple of sentences that momentum has been good lately, and you immediately get annoyed: could you not just show me the raw data table and the chart? It apologizes at once — sorry, I forgot that when you ask for a sales trend you want the table and the chart. In fact it did not forget; as a general-purpose agent, that is what it should do, unless you tell it to do it some other way. And you telling it that is you inventing a product ad hoc.

Third, its UI is designed for a general context, not a data context. For example, in Data Copilot a report may come to you as a card that you can open to see the data and then quickly return to the chat; in a general-purpose agent, a report is a link, and the link sends you off into another piece of software, another context. All of these are friction caused by product misalignment.

So, simply put: building a product means building a product. A product has to achieve alignment between cognitive form, brand, and function — I call this Product-Cognition Fit. Every point of misalignment generates friction: brand-cognition friction, where I have to talk myself into doing data analysis in a place called Code; context friction, where I ask it to find me a report and it finds me a pile of documents; UI friction, where I want an interactive chart and it returns a static png generated by Python matplotlib.

All of this friction is understandable and explainable technically and at the implementation level, but from the user's perspective it is all simply wrong — and users are under no obligation to listen to our explanations. They vote with their feet.

So what should we do? We should build a product, and achieve Product-Cognition Fit: use the name to carry an already-existing cognitive form (data work), carve out of the broad general context exactly the angle that matches data analysis, and then, in every detail of the product, execute it to match the user's expectations. When we get past the threshold, users will endure unimaginable inconvenience to use it.

What Exactly Did We Get Right?

A Simple Product Vision

What exactly did we get right, such that it became the data agent product people use — when those users have Claude Code and Codex?

First, a vision so obvious it is not worth mentioning: a chatbot with Data in its name; when users think of doing data work they open it, send a request, and get a result. Why is it not worth mentioning? Because when we say it out loud, there is nothing surprising about it. And precisely because of that, in its most basic form it matches the user's expectations exactly.

You might think that because I built an internal product, one major advantage is that users can always find the product's creator to solve their problems — but in fact, over the past six months, our users have very rarely come to me, including the power users who ask dozens of questions a day.

Which shows that our product form contains nothing users need to learn; they simply know how to use it. This is of course not a matter of intelligence, nor is it because our product execution is world-class; it is because we found the right thing and basically got it right. When a user uses a product, their mind is not blank. They arrive with a certain product imagination, and they know perfectly well: how complex is the thing I am doing, and therefore how much complexity should I expect from this product/tool.

I have seen too many complex BI products fail to work — not because their features are bad. They have all built plenty of features that are "powerful" in specific situations, but unfortunately that is not what users expect. When users have no choice, they compromise, and one signal of that compromise is that using the product becomes a skill. Take Tableau: you can see people write "proficient in Tableau" on a résumé. Why proficient? Because it is a skill. And when it is a skill, that means it does not match the average user's product imagination — because when it does match, it is not a skill.

Using Data Copilot is certainly not a skill. I can guarantee that not one of our users will ever write "can use Data Copilot" on a résumé — and that is the evidence that we matched the user's product imagination.

Second, the reason users do not come to me is that our product crossed the capability threshold. From interviews and surveys we can still see that users have plenty of opinions and even complaints, but they have worked around those problems. For instance, sometimes metric definitions are not perfectly aligned and the user has to guide it toward alignment — which is of course an excellent opportunity for us to make the product better, but from the standpoint of evaluating this product, it is also a signal of product completeness: the user can solve their problem inside the product, whereas in other products they might not be able to — traditional data and BI platforms, for example.

Of course, the technical reason users can solve their problems is that there is a general-purpose harness underneath (Codex). In principle, with enough effort, users could solve these problems in Claude Code and Codex too — but there is a cost question here: across a fairly wide range, Data Copilot happens to bring that cost down to the point where users can handle it themselves, and general-purpose agents do not. That is why, in the competition between general-purpose agents and Data Copilot, users still use Data Copilot.

Multi-Tenancy and the Workspace Distribution System

The origin of this product was our lightning realization that you could create a product by assembling the directory in which Claude Code/Codex does its work.

And once we turned it into a product, we faced a series of real problems:

First, it has to be multi-tenant.

Second, it has to have a centralized workspace template, and that template has to be pushable to every user — otherwise how could you call it a product?

Finally, users also need to customize their workspace, preserving the structures, conventions, and resources they prefer.

That is workspace management and distribution. We implemented a three-layer architecture:

  1. The centralized workspace template. This contains extensive resets of the general-purpose agent's "defaults," including output logic and format, the metadata protocol for output UI, data source selection and validation, and important SOPs and static infra knowledge. The directory where the end user works also needs various resources — code repositories, knowledge bases, document libraries — and we declare these in the template through a metadata protocol.
  2. The workspace base. By instantiating the template we get a base; the core of instantiation is assembling the resources declared in the template into the workspace. We generally maintain several bases at once, and users' workspaces inherit and pull from a base.
  3. The user workspace. This is what the user actually uses — the directory environment where Codex really works. Under certain conditions it periodically pulls updates from the latest base and merges them into its own workspace, so that the user's own overrides on top of the base are preserved. We also built concurrency control so that this synchronization does not interrupt the user's tasks.

To make this three-layer system work stably and smoothly, we had to design the workspace structure correctly from the start, along with the timing and mechanics of instantiation and distribution. Fortunately we got it right at the outset, and we have never once had a problem with it.

SEKE: A Self-Evolving Knowledge Engine

The workspace solves the question of where the product's defaults come from. But the genuinely hard part of data work is not in the defaults; it is in the organization's tacit knowledge.

The same metric is defined differently across business lines; a given table is only trustworthy after a certain point in time; a pipeline broke last week, so the data for those days has to be routed around; two user identifiers must be used in different scenarios, and sometimes one has to be derived from the other to do a join. None of this is in any model, and none of it is in any document (even though we think it should be) — it lives in the heads of the people who do the work.

A general-purpose agent cannot handle this layer, not because it is stupid, but because it starts from zero every time. You teach it once and it gets it right this time; the next time it is a different person, a different session, and it is wrong again. In the abstract, teaching it each time seems like a small thing, but this detail friction accumulates, and at some point the user stops coming back.

So we built a self-evolving knowledge engine (SEKE). The idea is simple: since this knowledge is spoken aloud over and over in the course of real work, capture it during real work, refine it, and let everyone share it.

Three things make it fundamentally different from "adding a memory to an agent":

First, knowledge is not memory. Memory grows linearly with experience — the more you do, the bigger and messier it gets. Knowledge grows through abstraction: ten corrections of the same kind should converge into one rule, not ten records. So the core of the engine is not storage but refinement: extracting reusable rules from concrete interactions, and revising old rules when new evidence appears.

Second, we do not want explicit user feedback. Asking users to thumbs-up, rate, or write comments does not hold up in a real work setting — they are doing their job, not training a model. So knowledge is inferred from the interaction itself: did the user accept this result, or change the metric definition and ask again, or take the answer and go use it? Behavior is far more reliable than stated opinion.

Third, the judge is reality, not the system. This is the point I consider most important. The easiest thing to challenge about a self-evolving system is whether errors will amplify themselves — it learns one thing wrong and then stays wrong forever. But in the data domain, the criterion is external: whether the numbers the SQL produces are correct, whether the report will be challenged on the spot when it is taken into a meeting, whether a wrong metric definition causes someone downstream to shout immediately. This domain comes with a built-in negative feedback loop, and that is exactly why it is suited to self-evolution.

At the same time, the knowledge tree cannot be handed over to the system entirely. We keep a layer of human-set structural constraints that the refinement process cannot override — which definitions are immovable, which boundaries cannot be crossed. The system can grow freely beneath that layer, but it cannot rewrite it.

The result is this: one person teaches it once, and the whole company is right from then on. Someone corrects a metric definition, and the next day a colleague in a different department doing a completely different analysis finds that the definition is already correct — without ever knowing this happened.

This is also the most fundamental difference from a general-purpose agent. A general-purpose agent's knowledge comes from training and belongs to the whole world; Data Copilot's knowledge comes from the work this group of people at this company did over these six months, and it belongs only here, and gets thicker every day.

I have written up the full design of this architecture in another post.

Why Domain Intelligence Exists, and the Agent as Interface

The above lays out why there should be a data chatbot (agent) product. Let me now talk somewhat more broadly about the question of domain intelligence.

What is domain intelligence? If there is a set of tasks that users have a concept for in their minds (for example "data analysis"), and a general-purpose agent cannot do it well out of the box, then that thing is called a domain, and the intelligence that does that task set well is called domain intelligence. If a domain is strong enough and frequent enough, then a product for that domain can hold up and can succeed.

We are in the agent era now, so people ask: could this capability be mounted onto general-purpose agents and general-purpose agent products?

First, from the standpoint of domain intelligence, by definition it must be built in isolation from general intelligence. The core issue here is not physical capability but the choice of angle: general intelligence is required and defined to handle general tasks from a general angle, so to give it a domain angle you must isolate it — that is, constrain it. For example, "find me a sales report" means different things in different contexts. The reason general intelligence cannot do this data analysis work well is precisely the reason it is general intelligence. Its inability to realize domain intelligence is therefore essential.

To realize domain intelligence you must constrain and isolate, and the final form of that constraint and isolation is a domain agent. In terms of product form, we might be able to mount a domain agent onto a general-purpose agent's UI, but you would still ultimately have to solve the context and UI customization problems discussed above. At the same time, you would have to pay this cognitive tax: the user goes to the general-purpose agent's UI and then has to choose which domain's work they want to do — this form might work, but it is genuinely a form of friction.

Another important issue is security and permission isolation. Some of the capabilities Data Copilot integrates could not possibly be opened up directly on everyone's laptop — the risk of doing so is unimaginably large. Integrating these capabilities and permissions into some relatively general internal agent might be possible, but that brings us back to the point above about domain intelligence requiring isolation.

So for as long as domains exist historically, the reasons for domain agents and domain agent products to exist probably still outweigh the reasons against. That is this author's judgment.

On "Wrappers"

Back to the "wrapper" objection.

When Manus first came out I studied it, and I was very surprised — I thought the level it reached was simply beyond most teams; it was certainly beyond me. Later, when Meta proposed acquiring it, someone scoffed: "What did you acquire? Isn't it just a few dozen gigabytes of Markdown?"

That statement is wrong, but you can understand from it what a product really is.

When an agent's underlying capabilities can be easily shaped and assembled with natural language, the essence of a product is exposed: what you have to build is a thing that carries the user's cognitive form. And the difficulty of that may be widely underestimated — because to this day, no other agent product has reached Manus's level in the areas Manus is good at, not even remotely close.

Which shows that building an agent wrapper is not that easy. In fact you can see how hard it is just by reasoning it through: you are competing with Claude Code and Codex, and for a user to choose your "wrapper" over Claude Code/Codex, you must have done something extraordinarily right — because the product teams behind Claude Code/Codex are the strongest in the world.

So where exactly is the difficulty?

I think first you have to identify, very precisely, one or a series of specific user product imaginations — that is, you have to understand very deeply what results and what process the user expects. Then you have to execute that product imagination well in every detail, so that the user's experience with your product crosses the threshold of "usable."

Both of these are extremely hard.

What's Next: A Data Platform Centered on the Data Agent

Before Data Copilot, the bottleneck in data work was query throughput. Most of the effort went into writing queries and getting them right. So you could not write very many queries, or you had to build a pile of infrastructure to avoid frequently writing large numbers of them.

Data Copilot has basically solved this bottleneck. When you need to establish something, the AI can almost always write the correct SQL. So you can use large numbers of queries to solve problems you used to solve some other way — or that you would never have bothered to solve at all.

But once the bottleneck moves, everything downstream has to be rebuilt.

With throughput up, the existing fixed servers cannot keep up and are not elastic enough, and splitting things out to run tends to contend for resources. So we built ClickHouse Serverless, using elastic instances to support high concurrency. Running dozens of queries in parallel and then aggregating a conclusion is now something we can do.

Going further: for complex requests, the agent can orchestrate and decompose them into a parallelized DAG, then dispatch that to serverless for execution.

In this way, the agent is no longer merely calling an existing capability; the capability itself has the agent deeply involved in it.

This is also why "adding a layer of agent on top of an existing engine" will not fully hold up — not because the layer is added badly, but because the original execution model was designed to receive one definite query, whereas what an agent produces is an intent that is still evolving. And once we can work at the level of intent, we can do many different things that nonetheless make sense.

Data Copilot is reshaping and penetrating into the underlying data infra. This is what we mean by a data platform centered on the data agent.

The Object of Governance Has Changed

Something else is happening at the same time.

Before the agent era, we did data governance mainly through "productizing reports": carefully maintaining and auditing a few dozen parameterized core reports, covering most needs through combinations of parameters, and then concentrating our effort on optimizing them. The control point was at the query layer.

In the agent era we still rely on these reports, but the core object of governance has become knowledge. As long as the knowledge is accurate, the agent can generate the correct query. The path is more direct, and the cost of bootstrapping is lower.

There is a real concession here: you give up strong control at the query layer.

I judge this concession to be worth it, on the grounds that it lets the organization develop naturally and encourages more data analysis and faster iteration. And of course, on the basis of the knowledge that accumulates, we could in the future have the agent automatically generate a set of core reports for humans to review.

But the direction seems clear to me: the knowledge that accumulates as users interact with the agent will become the primary object of data governance.

Finally

I built this product. Everything it can do could probably be done with Claude Code and Codex.

But today, for anything data-related, I still use Data Copilot almost exclusively.

The reason is simple: when I am the one using it, I do not want to invent a product.


I am writing a series on production AI agents — architecture, product judgment, and the problems you only run into once you have actually built one. If you are building something similar, you can subscribe here: blog.dreambubble.ai