Dev Blog

The Data That Won't Sit Still, Part 2: Teaching the Platform What the Data Means

Written by Funnel Dev | Oct 2, 2026, 9:35:05 AM

Written By: Anirudh Mehta, Engineering Manager

In Part 1, we wrote about the problem underneath marketing-data ingestion: yesterday’s data can change tomorrow. Keeping it correct means continuously deciding what to refresh, what to leave alone, and how to do that without re-fetching everything all the time.

But freshness is only one kind of correctness. You can ingest every row perfectly, store it in exactly the right shape, and still return the wrong answer. Ask something as simple as “what did we spend on Google Ads last month?” and the answer depends on more than the presence of a cost field. Which table should it come from? At what grain? Which date does “last month” refer to? What happens if the same fact appears in more than one report?

That’s the problem we’re building the new data platform to solve: adding meaning to what the data is, so that the platform knows how to give meaningful answers.

The problem isn't just ingestion: it's meaning

The other half is making sure that when someone asks "what did we spend on Google Ads last month?" then everyone, whether it’s the dashboard, the export, the analyst, or the AI agent, gets back the same, correct number.

That sounds simple. After a decade of connecting hundreds of platforms for thousands of customers, we've learned it isn't:

  • Grain is easy to get wrong. A campaign-level report and an ad-level report from the same platform describe overlapping populations at different levels of detail. Naively combining them can inflate or deflate a total, and the mistake doesn't announce itself; the number just looks slightly off.
  • The same fact can show up twice. A connector may expose the same metric through several overlapping reports: campaign-level cost, ad-level cost, device-level cost. Treating those reports as independent datasets and summing across them can silently double-count the same underlying activity.
  • A schema tells you what fields exist. It doesn’t tell you what they mean. A CRM deal has a created date, a stage-change date, and a closed date. A social post has a publish date and a lifetime of ongoing engagement. Collapsing all of that into one universal date field is a decision that quietly breaks certain reports.
  • Relating data across sources is hard without native joins. Revenue by lead source, budget pacing, product-level ROAS - these all require joining CRM, ad, and ecommerce data on keys that are sometimes messy, sometimes user-defined, and sometimes only valid for a window of time.
  • Our expression runtime is a product of its era. Our accumulated custom-field logic is powerful, but it also carries the quirks of the scripting language it was built on. Those quirks make it hard to port faithfully to a modern, vectorized execution engine.

None of these are exotic edge cases. They show up in ordinary reporting, and when they go wrong, the cost isn't abstract - it's a finance team that stops trusting the dashboard and goes back to spreadsheets, or a customer who spends a day rebuilding a CPC formula that the platform should have given them for free.

Building the new data platform: data and meaning, cleanly separated

Rather than patch these problems one at a time on top of the existing data platform, we're building a new layer that treats meaning as a first-class citizen, on the open-standards foundation we described in the last post.

The shape of it:

Connectors declare what they produce, not just how to fetch it. Instead of a connector being "a thing that downloads rows," it also declares which business concepts it provides, such as cost, impressions, campaign name, and which of its own tables relate to each other. That declaration is what everything downstream builds on.

Data lands in an open lake. As covered in Part 1, we're on Parquet with Iceberg table metadata, catalogued through Polaris. Each customer gets its own isolated warehouse. Nothing here is exotic anymore. It's the same open-table-format stack a growing share of the data industry has converged on, which is exactly the point: we're not the only ones who have to maintain it.

A semantic model sits between raw tables and a query. Above the physical tables, we compose a model from three ingredients: what the connector declared, an opinionated blueprint of which fields and relationships matter for a given use case, and anything the customer has customized.

Each of those layers contributes a different kind of knowledge. The connector knows the source: which fields it can produce, which tables belong together, what grain those tables represent, and which business concepts they expose. Funnel adds domain knowledge: which dimensions and metrics should be presented together, which relationships are valid, and how concepts such as cost, campaign, customer, or conversion should behave across different sources. The customer can then add the final layer of meaning, custom fields, mappings, relationships, or definitions that are specific to how their business works.

We compile those ingredients into a semantic model that becomes the surface consumers query. A dashboard, export, analyst, or AI agent asks for concepts such as “cost by campaign and date”; it should not need to know which connector table contains the metric, which physical grain is safe to use, or how that table relates to the rest of the model. Those decisions belong in the platform, where they can be made consistently and explained when they cannot be made safely.

The query layer resolves, compiles, and executes. Given a request and a compiled semantic model, the new data platform determines which fields, relationships, and physical tables can satisfy it, including which representation is valid when a metric exists at more than one grain. That semantic query is compiled into a Substrait plan using Semstrait, our open-source library for building semantic queries on top of Substrait. DataFusion then takes care of executing the resulting plan using Arrow’s vectorized execution underneath.

We’re intentionally not laying out the exact mechanics of how grain conflicts get resolved or how overlapping sources get de-duplicated. That’s the part of the engineering we’re still sharpening, and it’s also, frankly, some of the more interesting IP we’re building.. But the design goal is one we're happy to say out loud: wrong answers should fail loudly, not quietly. A platform that returns a number with no way to explain where it came from isn't actually solving the trust problem, it's hiding it.

Where things stand today

We didn't start by designing the whole platform, we started by proving the riskiest assumptions.

The first milestone was deliberately narrow: take existing custom-field logic, translate it into the new query representation, run it through the full new pipeline, from open lake to catalog to query engine to export, and compare the output against our current production system, row for row. It matched. More importantly, it proved that the underlying infrastructure, including the catalog, query engine cluster, and export path, is production-capable, not just a nice demo.

We're now working through the next slice: two live ad-platform connectors, a metric defined once and shared across both, and a semantic service that assembles the model automatically rather than us hand-wiring it. The target is simple to state and surprisingly informative to build: "total cost by date, across both platforms" - answered correctly, without touching our current engine at all.

What's next

Once that slice is solid, the next stretch of work is about proving the capabilities that our current engine genuinely struggles with:

  • Correct distinct counts - unique customers, unique campaigns - as a native aggregation, not a workaround.
  • Cross-source joins declared as first-class relationships, so spend-vs-budget or CRM-to-campaign questions don’t require customers to first move the data into another warehouse and maintain their own lookup tables.
  • Automatic protection against double-counting when overlapping reports carry the same metric, without asking the customer to manually exclude a source.
  • Getting a real customer onto the new platform for the simplest valuable slice we can ship honestly: connect, ingest, model, export - no dashboards yet, just clean data landing where they need it.

Each of those is a capability we can validate independently before we bet the wider product on it.

Why we're telling you this

The infrastructure underneath analytics is becoming increasingly standardized. Parquet, Iceberg, Arrow, and DataFusion are powerful building blocks. The harder problem for us is increasingly the layer above them: knowing what the data means well enough to know when a query is valid.

Storage and execution are becoming easier to standardize, but meaning is not. That’s the problem we’re building the new data platform to solve.