Most data stacks fall short in one very specific way.
Not because the tools are bad — but because they were never designed for the operator.
We have no shortage of great tooling:
- Replication: batch, micro-batch, streaming — from Fivetran to Airbyte and Confluent to Striim
- Orchestration: Apache Airflow, Dagster, …
- Transformation: dbt
- De-identification / PHI / PII handling: Private AI, Inpher, Protegrity
Individually, many of these are excellent.
The gap
As an operator, what you actually need is simple (but rarely available end-to-end):
- What datasets are flowing right now?
- Where are they breaking down?
- What depends on what?
- What happens if something fails at 2am?
- Who gets alerted — and what do they do next?
- How does one upstream change ripple through everything else?
Instead, you get:
- fragmented dashboards
- partial lineage
- logs scattered across systems
…and a lot of tribal knowledge holding it all together.
After a while, the system “works”… but only because a few people know how to keep it running. That doesn’t scale.
The glass pipeline
What’s missing is a control tower for data operations.
Done right, this creates what I think of as a “glass pipeline”:
- No black boxes.
- No guesswork.
- No digging through five systems to understand what’s happening.
This isn’t about replacing tools. It’s about making the ecosystem operable as a system.
The real question
At some point, the challenge stops being:
…and becomes:
Curious how others are thinking about this — are your data systems truly operable, or just… working?

