By Mick Hittesdorf, Senior Cloud Architect, KX
At last week's FISD conference in New York, I had the privilege of delivering a keynote address addressing one of the most urgent, yet overlooked, challenges in financial technology today: Is your market data actually AI-ready? And more importantly, can you prove it?
Every firm across capital markets is eagerly reaching for the promise of agentic AI, autonomous agents and machine learning models capable of reasoning, acting, and executing at speeds, scales, and precision that no human can match. However, there is a stark gap between promise and reality. When AI agents fail, it is rarely because the underlying Large Language Model (LLM) lacks intelligence; it is because the model was fed deficient, unverified, or flawed market data.
If you missed the keynote, here is a summary of why data quality must be mathematically proven rather than assumed, the unique ways AI fails when confronted with bad data, and how firms can measure and quantify their AI readiness.
In traditional market data operations, data quality is almost universally assumed. It’s live, it’s paid for, and vendors routinely claim their feeds are clean and complete. Yet, few firms demand objective, systematic proof.
Instead, firms rely on downstream discovery: traders, quants, and compliance analysts become the default alerting mechanism when a model makes a bad trade or a chart looks distorted. By the time a human spots the flaw, the damage is done, losses are incurred, regulatory reporting is compromised, and trust in the system is eroded.
Recent regulatory findings and academic research demonstrate just how pervasive and silent data defects really are across capital markets:
If human analysts struggle to spot these silent defects, autonomous AI systems stand no chance.
When bad data enters human workflows, experienced professionals often exercise intuition or spot obvious anomalies before taking action. AI, however, does not possess common-sense guardrails. AI will only amplify and magnify any defect that exists in your data.
Crucially, Machine Learning (ML) and Generative AI (GenAI) fail in fundamental, structurally distinct ways when supplied with deficient data:
To determine whether market data is ready to feed into automated models, firms must evaluate their data pipelines across six core questions that AI cannot answer on its own:
To illustrate how missing context breaks AI models, consider Walmart’s (WMT) 3-for-1 stock split on February 26, 2024.
On that day, shares tripled and the share price was divided by three. Economically, nothing changed about the company. However, if unadjusted market data is fed to a model, it registers an abrupt 67% price crash. An agent processing this unadjusted data sees a catastrophic drop (Accuracy/Completeness test failure = Low AI Readiness Score). A model processing corporate-action-adjusted data recognizes the split seamlessly (High AI Readiness Score).
Data is only AI-ready when systematic, automated, and continuous testing proves it is fit for its intended purpose at a specific point in time.
Firms must move away from static data quality claims and adopt an automated testing methodology:
We define the AI Readiness Score as a clear, quantifiable metric evaluated at a specific point in time. By setting strict score thresholds, firms can automatically prevent unverified or failing datasets from reaching live trading models, execution algorithms, or GenAI agents.
As financial institutions race to deploy autonomous workflows and predictive models, the key differentiator will not be who has access to the newest LLM, but instead who has built the most trustworthy, verifiable data foundation.
If you are not sure whether your market data pipelines can withstand the scrutiny of AI, or if you want to learn how to implement continuous AI readiness testing within your organization, I would love to connect.
Let's Talk:
Want to learn more? If you’re building systems that need market data infrastructure to power your AI initiatives, request a demo today.
Best wishes,
Mick Hittesdorf