Data platforms

Your AI is only as good as your data platform.

When an AI pilot disappoints, the model is rarely the problem. Almost always, the agent couldn't reach the right data, the data disagreed with itself, or nobody could say which number was correct. Here are the five foundations we check before we build anything, and what to do when they're missing.

1. Reachable

The question: Can an agent get to the data it needs through a governed interface, or is it locked in spreadsheets, email attachments, and a database only one person can query?

Symptoms when missing: Every AI project starts with weeks of "getting access." Extracts are emailed around. The agent works in the demo on a hand-built file and fails in production.

The fix: Land the core sources in a central platform, such as a lakehouse on S3 or a warehouse like Snowflake or Redshift, with automated ingestion. Expose them through a semantic layer or Model Context Protocol servers so agents and people use the same doors.

2. Trustworthy

The question: If two systems disagree about a customer's balance, which one is right, and does the platform know?

Symptoms when missing: Two dashboards, two numbers. Analysts spend their time reconciling. An agent picks one at random and confidently reports it.

The fix: Define ownership per domain, implement data quality checks in the pipeline (freshness, completeness, referential integrity, business rules), and fail loudly when they break. Tools like dbt tests and pipeline-level assertions make this routine.

An agent will act on bad data with exactly as much confidence as on good data. Quality checks are not a nice-to-have; they are the agent's immune system.

3. Defined

The question: Does "active customer" mean the same thing in every report and to every agent?

Symptoms when missing: Metric drift. The AI assistant gives a revenue number that doesn't match the board deck, and trust collapses.

The fix: A semantic layer: metrics and dimensions defined once, in code, and consumed everywhere. This is the single highest-leverage investment for analytics agents, because it lets the model translate a question into governed definitions rather than guessing at SQL.

4. Fresh

The question: How old is the data when the agent reads it, and does that match the decision being made?

Symptoms when missing: Nightly batches feeding a support agent that answers questions about today's orders. Or the opposite: expensive streaming for data nobody needs faster than daily.

The fix: Match latency to the use case. Most agent workloads are fine with hourly or daily batch through Airflow or Step Functions. Where minutes matter, add Kafka or Kinesis for those specific streams. Don't pay for real-time everywhere.

5. Permissioned

The question: When an agent answers a question for a user, does it only see what that user is allowed to see?

Symptoms when missing: The pilot runs with an admin credential. Security halts the rollout. Rightly.

The fix: Row- and column-level security in the platform, and agents that act on behalf of the user rather than as a super-user. Put the enforcement in the data layer and the tool layer, not in the prompt.

The readiness scorecard we use

Score each foundation from one to five for the specific data your first agent needs. Anything under three is a prerequisite, not a follow-up. In our experience, a focused four-to-eight-week effort on the gaps almost always pays back faster than a stalled AI program.

Fixing gaps without a multi-year program

The classic mistake is to conclude "we need a data transformation" and launch an eighteen-month platform initiative. You don't. Scope the foundation to the first agent's needs: two or three sources, one domain, one semantic model, quality checks on the fields that matter. Build that in weeks, ship the agent, and extend the platform as the next use case demands. The agent pays for the platform, and the platform makes the next agent cheap.

This is why our data platform and AI practices are one team. We check the foundations in the readiness sprint and fix what's needed as part of the pilot, so the AI investment lands on solid ground.

Not sure your data is ready?

The readiness sprint scores these five foundations against your first use case.