AI data readiness · Governance · Infrastructure

Most AI initiatives fail on the data, not the model.

Your model is fine. Your retrieval architecture is fine. What breaks is undocumented tables, conflicting definitions, missing lineage, and permissions that quietly stop applying three hops downstream. SchemaVita measures whether your data estate can actually support what you're building — and then fixes what can't.

The problem

The failure mode is upstream of the model.

This is not a hunch. It is the most consistently reported finding in enterprise AI deployment, and it is measurable before you spend another quarter on a pilot.

60%
of AI projects will be abandoned by organisations that lack AI-ready data.
Gartner projection, through 2026
52%
of questions produced fabricated answers when the same RAG system ran on unvetted data. On curated content, hallucinations fell to near zero.
Published medical RAG study
+9.2
percentage-point gain in retrieval precision from metadata enrichment alone — with no change to the retrieval architecture.
Published RAG benchmark
Services

Three ways to work together.

Start with the assessment. It is a real deliverable, priced on its own merits, and it is never credited back against later work — that would make it a sales call rather than an audit.

AI Data Readiness Assessment

Fixed fee · 1–5 weeks

A measured verdict on whether your data can carry the AI system you intend to build. Seven dimensions, scored on evidence rather than interviews, delivered as a scorecard, a findings register, and a dependency-ordered remediation roadmap your team can start on Monday.

  • Instrumented measurement of quality, coverage, lineage, entitlement, structure, and freshness
  • Scored against published standards — DAMA-DMBOK, NIST AI RMF, NIST SP 800-53
  • Findings tied to your specific use case, whether that is RAG, agentic, or predictive
  • Sequenced remediation roadmap with owners, effort, and prerequisites
  • All measurement artifacts handed over so you can re-run the numbers yourself
Focused
$5,500
1 week · ~12-page report
One data domain or a single AI use case. Two to three interviews. For teams with a specific decision pending.
Seed stage · small ops team
Standard
$9,500
2–3 weeks · ~30-page report
Full data estate, up to four source systems, five to eight interviews. The right scope for most companies with a stalled or pre-launch AI initiative.
Series A–B · mid-market
Deep
$18,000+
4–5 weeks · ~60-page report
Multi-domain, regulated, or multi-entity estates. Twelve to eighteen interviews, compliance mapping included.
Regulated · PE portfolio

Implementation

Scoped from the roadmap · $40k–$120k typical

Building what the assessment says needs building. Scoped as a fixed-fee project against the remediation roadmap, so you are buying a defined outcome rather than an open-ended engagement.

  • Data governance for AI — documentation standards, quality contracts, ownership, lineage, and the monitoring that keeps them true
  • Governed AI-assisted development — rolling out Claude Code, Cursor, or Copilot with enforced standards, hooks, security guardrails, and review gates
  • Agentic and self-healing pipelines — pipelines that detect their own failures, quarantine bad data, and recover without a human
  • Warehouse and transformation modelling — canonical entities, stable keys, and semantic consistency, with AI-readiness built in from the start

Fractional AI Infrastructure Engineer

Retainer · from $5,000/month

A fixed number of days each month for teams that have the systems but not the specialist. Ongoing governance ownership, standards enforcement, pipeline reliability, and someone accountable for the thing you just built continuing to work.

  • Set days per month, agreed in advance
  • Continuous coverage and quality monitoring, not a report once a quarter
  • Design review for new AI features before they ship
  • Typically follows an implementation; occasionally stands alone
Method

Seven dimensions, measured against published standards.

The framework is derived from public sources rather than invented, so you can check the work. Each dimension is scored on a five-level maturity scale, weighted for AI readiness specifically rather than general data health.

D1

Data Quality

Accuracy, completeness, consistency, timeliness, validity, and uniqueness — profiled and measured, not asserted. Completeness and consistency carry extra weight, because missing or conflicting data is what produces unreliable models.

DAMA-DMBOK six-dimension set
D2

Metadata & Documentation

Whether your estate is legible to a retrieval system. A table of accurate numbers is not AI-ready if nothing can discover it, its columns are cryptic, or its provenance is unknown. Coverage measured at model and column level.

dbt project evaluator · dbt-coverage · catalog metrics
D3

Lineage & Traceability

Whether any value can be traced to its origin and forward to its consumers. Required for reproducing model results, for staleness detection, and for GDPR, HIPAA, and EU AI Act obligations.

OpenLineage · catalog-native lineage
D4

Access, Entitlement & Sensitivity

Whether governance controls survive the trip downstream. Sensitivity classification has to happen before indexing — repairing access control at the vector store is already too late.

NIST SP 800-53 · SOC 2 Trust Services Criteria
D5

Structure & Semantics

Whether the same concept looks the same everywhere. Canonical fields, stable identifiers, normalised units and time zones, and one agreed definition per metric across teams.

Published AI-ready data literature
D6

Pipeline Reliability & Freshness

Whether the data arrives, on time, and whether anyone finds out when it does not. Freshness, test pass rates, alerting coverage, and how often a human has to intervene by hand.

Source freshness and run artifacts
D7

AI Governance Posture

Cross-cutting. System inventory, ownership, risk categorisation, evaluation practice, incident response for model failure, and third-party model visibility.

NIST AI RMF 1.0 — Govern, Map, Measure, Manage
How it runs

Measure first. Interview second.

Most assessments are interviews with a scorecard attached. This one instruments your estate before anyone gets asked a question, so the conversations are about interpreting evidence rather than collecting opinions.

01

Scope

A short call to fix the AI use case in question. RAG, agentic, and predictive systems have different failure modes, and the weighting changes accordingly.

02

Measure

Scripted instrumentation against your warehouse and transformation layer. Coverage, quality profiling, lineage completeness, entitlement inheritance, freshness.

03

Interpret

Interviews with owners, consumers, and whoever is driving the initiative — to explain the numbers and to score governance and entitlement, which artifacts cannot show.

04

Deliver

Scorecard, findings, and a dependency-ordered roadmap, walked through with your decision-maker. No pitch in that meeting.

About

Built by someone who does this at scale.

SchemaVita is Cordero Perez. I build the governance layer that makes AI systems accurate, reliable, and safe to deploy — documentation quality systems, entitlement monitoring, automated security guardrails feeding into AI build workflows, and the standards infrastructure that lets agentic systems act on organisational data without producing confident nonsense.

I do that work today inside a Fortune 500 technology company's finance organisation, under real compliance pressure. Before that, three years as a senior consultant in AI and data engineering at a Big Four firm, and three years as a data analyst for a municipal oversight and investigations division, where the work was cited in the New York Times.

The name means roughly structure brought to life — which is the job. Design the thing properly, then make it run.

Data & Analytics Developer
Fortune 500 technology · Finance · current
Senior Consultant, AI & Data Engineering
Big Four consultancy · 3 years
Data Analyst, Oversight & Investigations
Municipal government · 3 years
Certifications
AWS Cloud Practitioner · MIT/edX Supply Chain Analytics · MIT/edX Supply Chain Technology & Systems · Tableau Desktop Specialist · PCEP Python · SOA Exam P
Get in touch

Find out whether your data can carry it.

Tell me what you are trying to build and where it is stuck. If an assessment is not the right thing, I will say so — sometimes the answer is one conversation, not an engagement.

cperez@schemavita.com