A pipeline can run green every night, every dashboard can refresh on schedule, and the numbers inside can still be wrong. Nothing crashes.
No error fires. A source system silently starts sending null values in a column that used to be populated, a schema change upstream drops three fields nobody flagged, or a batch job that normally loads twelve million rows loads four thousand — and every downstream job still finishes "successfully," because success was only ever defined as the script ran, not the data is correct.
That gap between a pipeline succeeding and the data being right is the specific problem data observability tools exist to close. The term traces to a specific person and date: Barr Moses published the framework in December 2020, at Monte Carlo, modeling it directly on DevOps observability's metrics/logs/traces triad and translating it into five data-specific signals.
Five years later it's a real tooling category with real vendors, real acquisitions, and, as of 2025-2026, a genuinely unsettled market worth understanding before picking a tool. This piece covers what the category actually monitors, how ML-based platforms differ from rule-based testing frameworks, where the landscape stands after three separate restructuring events in eighteen months, and how to build toward this without overcommitting to one vendor too early.
Prefer to skip the debugging and have an expert handle this? Book a 30-min diagnostic →
What Data Observability Actually Monitors: The Five Pillars
Barr Moses's original framework defines data observability as an organization's ability to fully understand the health of the data in its systems. IBM's own definition arrives at the same substance independently: monitoring, managing, and maintaining data so its quality, availability, and reliability hold up across every process, system, and pipeline it passes through. The five pillars themselves are freshness (is the data current, and are there gaps in when it last updated), volume (is the amount of data landing within expected bounds, often the fastest signal that an upstream job silently failed), schema (are fields being added, removed, or retyped without anyone being told), distribution (are the values inside a column still statistically normal — null rates, ranges, cardinality), and lineage (the map of upstream sources and downstream consumers, used to trace blast radius once something breaks).
What both definitions agree on, and what separates this category from older data-quality practice, is automation: observability tools are built to catch problems nobody wrote a rule for in advance, not to check a fixed list of known invariants. That's also the category's real limit. All five pillars monitor structural and statistical properties — a table can pass freshness, volume, schema, distribution, and lineage checks simultaneously and still contain values that are wrong relative to what the business actually means by correct, because none of the five pillars encode business logic.
Evaluate a platform on how much table coverage it reaches with zero configuration and how it behaves in its first weeks after connecting, not on pillar coverage alone — every vendor in this space now claims all five.
ML Anomaly Detection vs. Rule-Based Testing
Two genuinely different mechanisms get sold under the same data observability label, and knowing which one you're buying matters. ML-based platforms — Monte Carlo, Bigeye, Acceldata, Anomalo — infer what normal looks like from historical data and alert on statistical deviation without anyone writing the rule in advance. Rule-based frameworks — dbt tests, Great Expectations, Soda's SodaCL — require an explicit, human-authored assertion; dbt's own four built-in generic tests are unique, not_null, accepted_values, and relationships, extendable with custom SQL.
The two aren't really competitors, they're different coverage strategies. ML-based monitoring earns its keep when the table count is too large to hand-write rules for all of it — Bigeye's own docs describe a library of more than 70 pre-built data-quality metrics precisely because writing dozens of rules per table doesn't scale past a few hundred tables. Rule-based checks earn their keep on known, business-critical invariants — primary keys, referential integrity, a specific business rule — where an explicit, reviewable assertion is worth more than an inferred one, because a reviewer can read exactly what it checks.
# dbt: an explicit, human-authored rule
models:
- name: orders
columns:
- name: order_id
tests: [unique, not_null]
- name: customer_id
tests:
- relationships:
to: ref('customers')
field: customer_id
# ML-based observability: no rule written —
# the platform learns "normal" row count, null rate,
# and schema shape from 3-6 weeks of history insteadThe Cold-Start Problem Nobody's Marketing Page Mentions
There's a real gap in ML-based monitoring that neither vendor foregrounds in its marketing, though both disclose it in their own technical docs. Monte Carlo's documentation states thresholds typically start generating after seven days and mature by fourteen, trained on a rolling six-week window. Bigeye's documentation defaults to a twenty-one-day training window and states reliable seasonality detection needs three or more cycles of a pattern — at least three weeks of history for anything on a weekly cadence.
Put plainly: a newly onboarded table, or an existing table whose pattern just changed, has the weakest ML-based protection during exactly the window it's statistically most likely to break. That's not a vendor flaw you can route around by switching platforms — every agentless ML monitor shares the same structural cold-start, because none of them can infer a pattern from history that doesn't exist yet. Rule-based checks don't have this problem: a not_null test is exactly as strict on day one as on day one hundred.
It's also worth watching what happens after the training window closes. Monte Carlo's own documentation describes automatically widening a threshold when the same type of anomaly recurs often enough — genuinely useful for cutting alert fatigue on benign recurring patterns, but also a mechanism that can quietly absorb a real, recurring regression into the new normal if nobody periodically reviews what the system has auto-tolerated. Neither vendor documents a built-in audit trail for reviewing those auto-widened thresholds as a standing practice — that's a process a team has to build on its own, not something the tooling hands you by default.
Agentless Monitoring vs. Pipeline-Embedded Checks
The other structural split worth understanding before buying anything is where the check actually runs. Agentless platforms — the dominant architecture at Monte Carlo, Bigeye, Metaplane, and Acceldata — connect read-only to a warehouse or lake and infer health from metadata and query logs after the data has already landed. Pipeline-embedded checks run inside the pipeline itself, before the data is considered done — a failing dbt data test fails the dbt run outright, and Dagster's asset checks can block a downstream asset from materializing at all when a check fails.
This isn't a vendor-quality difference, it's architectural: by construction, an agentless platform only detects a problem after the data has already landed and potentially been read by a dashboard, a report, or a model, because it's observing the warehouse rather than gating the pipeline that fills it. Agentless monitoring is the right call for retrofitting broad coverage onto an existing warehouse fast, without touching pipeline code. Pipeline-embedded checks are the right call anywhere the requirement is don't let bad data get consumed at all, not just tell me quickly after it happened.
The practical rule: let a table's business-criticality decide which pattern protects it, not its size. Put blocking, pipeline-embedded checks on the highest-blast-radius tables specifically, and reserve broad agentless ML coverage for the long tail of tables nobody's gotten around to writing explicit rules for yet.
Still Piecing This Together Yourself?
A senior engineer looks at your actual setup, not a generic checklist, and tells you exactly what's wrong and how to fix it.
Book a Diagnostic CallThe Vendor Landscape, and Why It's Mid-Restructuring Right Now
The landscape splits roughly into agentless ML platforms and open-core rule-based frameworks, and most real deployments end up running one of each. Monte Carlo, the platform that coined the term, has repositioned itself around agent observability alongside its original five-pillars product — a genuine strategic shift, worth asking about directly if core data observability is what you actually need. Bigeye, a 2021 rebrand of a product originally called Toro, documents its anomaly-detection method more explicitly than most competitors — a five-step statistical pipeline covering structure analysis, blind prediction testing, uncertainty modeling, boundary calculation, and sensitivity adjustment — and, like Monte Carlo, launched its own AI-agent-focused product line in mid-2025.
Datadog acquired the independent startup Metaplane in April 2025 and folded it into a broader Data Observability suite alongside Data Jobs Monitoring and Lineage — the one platform here that can correlate a data-quality alert with the infrastructure metric that caused it, if you're already a Datadog customer. IBM's Data Observability by Databand sits inside the watsonx.data line and shows a materially slower release cadence than its independent competitors. Acceldata sells the broadest, least-focused surface of the group: observability plus agentic data management plus its own warehouse product, aimed at regulated and hybrid-infrastructure buyers.
On the rule-based side, Soda ships a genuinely free, Apache-licensed core with a commercial collaboration layer on top; Great Expectations went through a real split in mid-2026, when FICO acquired its commercial GX Cloud product and discontinued it while Fivetran took over stewardship of the still-free GX Core; and Elementary built the deepest dbt-native integration in the category, storing checks as version-controlled dbt artifacts rather than a separate config surface. Anomalo and Datafold round out the field, and both have drifted furthest from classic five-pillars framing — Anomalo toward a suite of specialized AI agents for table monitoring, Datafold toward AI-assisted migration tooling with anomaly detection as a supporting feature rather than the product's center. Three independent ownership or positioning changes inside eighteen months is not a settled market.
It's a category actively being redrawn, and that's worth factoring into any multi-year vendor commitment.
Lineage and the Standards Layer: OpenLineage and Data Contracts
Two open standards sit underneath this landscape, and neither is vendor-specific — which matters given the ownership churn above. OpenLineage, governed as a Linux Foundation graduate project, is an open framework for collecting and analyzing data lineage: a core model of three entities (Dataset, Job, Run) extended through facets, user-defined metadata attachable to any entity without changing the base spec. One facet is directly relevant here — dataQualityAssertions — which captures the result of a data test (a type like not_null or unique, a pass/fail boolean, an optional severity distinguishing a blocking error from a non-blocking warning) as portable lineage metadata rather than something that only lives inside the tool that ran the test.
OpenLineage ships as a native, first-party integration inside Airflow, Spark, dbt, and Great Expectations. Whether the ML-based observability vendors specifically emit or consume that facet is genuinely unverified — worth asking directly in any vendor evaluation, not assuming. The second standard, the Open Data Contract Standard (ODCS, governed by Bitol under the Linux Foundation), formalizes the upfront agreement between a data producer and consumer team as a vendor-neutral YAML schema, with a data-quality rules section that can reference Soda, Great Expectations, dbt, or Monte Carlo as the actual enforcement engine.
ODCS's own documentation is candid that this vendor-specific wrapping is an intermediate step, not yet native cross-vendor portability — the contract layer is standardized today, the detection layer underneath it still isn't.
What This Actually Costs
Pricing transparency is itself a real differentiator in this category, and it splits sharply. Monte Carlo publishes an actual number — its Scale-tier order form states a rate of $0.28 per credit — but the number of credits a given monitor configuration actually consumes is incorporated by reference to a separate consumption-rate document, not disclosed in the contract itself, which means a buyer can see the unit price without being able to fully project total spend from the order form alone. Metaplane, now under Datadog, prices per monitored table instead, which inverts the incentive: broader coverage doesn't cost more per credit consumed, it just costs more tables.
Every other vendor covered here — Bigeye, Acceldata, IBM Databand, Anomalo, Datafold — publishes no pricing at all; all five run a sales-led, custom-quote process, confirmed independently across each vendor's own site, not a coincidence but a category-wide norm at the enterprise tier. Open-source doesn't fully dodge this either: Great Expectations' commercial cloud layer no longer exists after the FICO acquisition, and Soda's own documentation states plainly that the free Soda Core tier lacks observability features — the free layer is real, but deliberately capability-limited by design, not a scaled-down version of the paid product. The practical implication: request each vendor's actual per-monitor or per-credit consumption-rate table before signing anything, not just the headline number — the headline number alone doesn't let you project a real bill.
Get a Straight Answer, Not a Sales Pitch
Tell us what you're running into. We'll tell you directly what's causing it and what it takes to fix, before you sign anything.
Contact UsHow to Build Toward This Without Overcommitting to One Vendor
Given the pricing opacity and the pace of restructuring above, the more defensible starting point for most teams isn't pick a platform, it's build the open-source-anchored foundation first and treat any single commercial platform as a replaceable layer on top of it, not the system of record. That foundation is three things: dbt tests on every primary key at minimum, OpenLineage emission wired into whatever lineage backend you land on, and a free-tier observability layer — Elementary OSS or Soda Core — covering the tables that don't yet have explicit rules. Layer pipeline-embedded blocking checks specifically onto your highest-blast-radius tables, where the cost of bad data reaching a dashboard or a model is genuinely high, and reserve broad agentless ML monitoring for the long tail of tables nobody's gotten around to writing rules for.
Track the category's own self-defined success metric rather than a generic fewer-incidents number: data downtime, defined by Monte Carlo as the number of incidents multiplied by the average time to detect plus the average time to resolve — a direct analogue to DevOps' MTTA/MTTR discipline, and one that's genuinely comparable before and after a change, regardless of which vendor ends up in your stack. Only add a commercial ML platform once you know your real table count, your actual budget, and which of the two coverage models your highest-risk tables specifically need. Picking a platform before that point means betting on today's category leader continuing to look the same way in eighteen months, and the last eighteen months say that's not a safe assumption.
Related: Data Reliability
Data observability isn't a single product, it's two complementary mechanisms — ML-based anomaly detection for the unknown-unknowns, rule-based testing for the known invariants — wearing one marketing label. The five pillars are now table stakes across every vendor, so the real evaluation criteria are the ones vendors don't lead with: how long a platform takes to trust a new table, whether a check blocks bad data or just reports it after the fact, and how exposed you are if the vendor you pick gets acquired, re-priced, or repositioned before your contract renews. Given three restructuring events in eighteen months across this category, the safer sequence is foundation first — dbt tests, OpenLineage, a free observability layer — commercial platform second, sized to a table count and budget you actually know rather than one a sales call talked you into.
Frequently Asked Questions
What is data observability?
Data observability is the automated, typically ML-driven practice of continuously monitoring data pipelines and warehouses for reliability problems nobody explicitly wrote a rule for — commonly measured across five properties: freshness, volume, schema, distribution, and lineage. Barr Moses defined the term at Monte Carlo in December 2020, modeling it directly on DevOps observability's metrics/logs/traces framework.
What are the five pillars of data observability?
Freshness (is the data current), volume (is the amount of data within expected bounds), schema (are fields changing unexpectedly), distribution (are column values still statistically normal), and lineage (the upstream-source/downstream-consumer map used for impact analysis). Nearly every vendor in the category now covers all five; the real differentiator is how each detects and prioritizes anomalies against them.
What's the difference between data observability and data quality testing?
Data quality testing (dbt tests, Great Expectations, SodaCL) requires someone to write an explicit rule in advance and typically runs inside the pipeline, blocking bad data before it's consumed. Data observability (Monte Carlo, Bigeye, Acceldata) uses ML to infer normal from historical data and catch problems nobody wrote a rule for, but usually only after the data has already landed in the warehouse.
Is Great Expectations still free?
The open-source core, GX Core, remains free and Apache-2.0 licensed under Fivetran's stewardship as of mid-2026. The commercial SaaS layer, GX Cloud, was acquired by FICO and discontinued on June 1, 2026 — teams currently on GX Cloud need a concrete migration plan.
How much do data observability tools cost?
Pricing is opaque across most of the category. Monte Carlo publishes a $0.28-per-credit rate but not the actual consumption rate per monitor; Metaplane, now under Datadog, prices per monitored table; Bigeye, Acceldata, IBM Databand, Anomalo, and Datafold all use sales-led, custom-quote pricing with no public numbers. Request a vendor's actual consumption-rate table before contracting, not just the headline price.
Do I need a dedicated observability platform, or are dbt tests enough?
dbt tests catch known, explicitly defined problems and block bad data before it's consumed, but only for what someone thought to test. A dedicated ML-based platform catches unanticipated anomalies across an entire warehouse, at the cost of a cold-start window before it's fully trained. Most mature setups use both: pipeline-embedded tests on the highest-risk tables, ML monitoring on the rest.
What is "data downtime"?
Data downtime is the category's own self-defined success metric, coined by Monte Carlo: the number of incidents multiplied by the average time to detect plus the average time to resolve. It's a direct analogue to DevOps' MTTA/MTTR, and it's a more useful before/after measure of an observability program than raw incident count alone.
Who can help me set up data observability without picking the wrong platform?
DharmOps builds the open-source-anchored foundation this guide recommends — dbt tests, OpenLineage, a free-tier observability layer — under Data Reliability, before recommending any commercial platform, sized to your actual table count and budget.
How much does it cost to implement data observability?
The tooling itself ranges from free (dbt tests, Soda Core, Elementary OSS) to sales-quoted enterprise platforms with opaque per-credit or per-table pricing, as this guide details. The engineering cost of wiring it correctly is the more predictable line item — worth scoping with a Diagnostic Call before a vendor sales call sets the number.
Do I need a consultant to set up dbt tests and OpenLineage, or can my team do it?
A team already running dbt can usually add tests and OpenLineage emission directly. Where outside help earns its cost is deciding which tables get blocking, pipeline-embedded checks versus broad agentless ML coverage — the tradeoff this guide walks through — since getting that split wrong either misses real incidents or burns budget on low-risk tables.
Get a Straight Answer on Your Setup
Tell us what you're running into. We read every message personally and reply within 24 hours with times for a free call.

