By Peter Simpson, OneTick Product Owner, KX
Following our recent workshop on accessing on-demand market data from KDB-X, I had the opportunity to answer a series of audience questions. While the session was designed for KDB financial professionals interested in simplifying their market data pipelines for quantitative development within KDB-X, the questions we covered apply to users of all analytics platforms.
OneTick Cloud is our high-quality, on-demand managed time-series data and analytics platform that provides users with global, AI-ready market data, instantly seamlessly fueling compute, research, and analytics engines.
Data hydration is the process of taking raw, fragmented, or asynchronous financial data and standardizing, enriching, and cleaning it so it becomes AI-ready. It bridges the gap between raw unstructured data (like raw exchange ticks) and actionable machine learning and quantitative analysis.
The hydration process involves several steps:
For algorithmic trading and AI agents, hydrated data is crucial because it eliminates the "data tax"—the 70% to 80% of time quants typically spend cleaning feeds. By providing fully harmonized, audit-ready time-series data, it allows trading systems to operate without the data gaps that cause hallucinations or inaccurate predictions.
With that context established, let’s get into the questions.
The installation of the module probably takes ten to thirty seconds by the time you've downloaded it and copied it into your KDB-X install.
That comes with prebuilt credentials. So from start to finish, in terms of accessing the samples, that takes thirty seconds. In terms of accessing more datasets, which could be any market globally, whether that's equities, futures, or options, Once the paperwork is signed, we set you up, which is just giving you your own credentials, and then you have access to the full history.
So once paperwork's signed, that's probably an hour or so before you can get your new credentials.
It takes a bit longer in terms of real time because then we also have exchange agreements to be signed.
But it's literally hours. There is no backloading of data, waiting for data to be loaded onto your local storage, providing servers. It's just, “here's the data.” You have pretty much immediate access to it.
Okay. Two main ways. The first way is you just go to Google, type "OneTick SQL documentation," and you'll come into our public, Sphinx based documentation.
And that has an AI assistant built in. So you can either read through the documentation or just ask a question, and that will come up with the syntax for you.
Or you can, use our MCP server, connect up, and then ask through your IDE of choice, let's say, Versus Code, or just kind of Claude code, ask the natural language question, and that will come back with the SQL you should use to execute.
If you go to the OneTick website and click on Market Data and then Coverage, you can see all of the datasets that we provide. That's also linked from the KX website. And there are details to drill in and then see every single schema that we provide. So under underlying each dataset will have a series of databases, which will have a series of tables, either the kind of underlying trades quotes, MBBO, book depth, but also derived data sets like one minute trade in quote bars, and also day summaries where we roll up and kind of which you'll see in split volume by trade types, and then also add in the kind of feature sets such as kind of enriched trades. So it's easy to go to the website and just look.
Also, all this information is available on GitHub. If you go to the one market data GitHub, you can pull back the information there. And also through our MCP server, you can ask a natural language question, like, which database holds Italian equities and come back with the result, which will be Milan, but also there are other datasets, kind of pan European datasets that also hold Italian equities. Then we also have our European composite, which aggregates all of that flow across each venue.
So lots of ways either using our MCP server or using our websites or GitHub.
Typically on a monthly basis, where we're adding new venues, that is less now typically around new venues and more around adding book depth for an existing venue. There are also regulatory changes.
So for example, for the US SIP, we've added additional tables for Odd Lot quotes and Odd Lot quotes in part of the NBBO that happened in June.
We're also continuously adding new AI feature sets. So last month, we added for the US enriched trade table.
So rather than just putting a table that includes every trade historically, this table includes every trade and additionally, the prevailing quote from the NBBO and also markouts based on the NBBO and around thirty markouts ranging from minutes before two minutes after the NBBO.
Just to make your analytical questions easier, rather than having to do a whole load of compute yourself, you can just go to the AI feature set and pull back the dataset that's enriched for your purpose.
There are a few different aspects here:
Are you interested in accessing a whole venue? So do you want every single US equity, for example, or European equity? Or are you more interested in just, let's say, the top six hundred equities or the top one hundred futures products, you tell us which you want. We'll provide the most effective option.
Do you want access to the whole venue, or do you want to be more so venue priced or symbol priced? And then how much history do you want to receive, ranging from one year of history going back to to twenty years.
Do you require access to real time data, or are you just looking at T+1?
And then how much compute do you require to run your queries? Are you just dumping data down, or do you want to run the SQL to aggregate the data and just pull back the summary data?
We'll work with you based on your requirements. Our existing customers typically start subscribing to a small range of data, and over time, they get more comfortable and expand their coverage universe, either adding new markets or increasing the history that they subscribe to.
Please try it out for yourself. As we saw from the demo, it's a thirty second install, issue the SQL query, and come back with the data. The documentation's on the KX website. Ask us if you have any questions.
OneTick Cloud offers a single vendor end-to-end solution, significantly reducing the time to value. KX and OneTick are the only vendor delivering AI-ready, hydrated, temporal market data as a managed service. Our data is pre-normalized across 250+ venues and 30+ years of history, point-in-time with no look-ahead bias, machine-readable from day one, and fed natively into Python, SQL, and KDB-X.
Save time, money, and resources by letting the KX OneTick team clean feeds, map symbols, and align timestamps so your quants, analysts, and AI models can do their work.
Or you can always schedule a call with the OneTick team today: contact us here.
Best wishes,
Peter Simpson