Custom datasetthe catalog doesn't have it?
the catalog doesn't have it?
we'll scope it.
Tell us what you need and how you intend to use it. We'll come back with scope, documentation and licensing — and we'll tell you early if it's the wrong tool for the job.
brief → feasibility → quote → generation → verification → delivery
Custom flow
six steps, same dossier standard.
A custom dataset goes through the same verification as a catalog dataset — just built for your scenario instead of someone else's.
- 01Brief
- 02Feasibility
- 03Quote
- 04Generation
- 05Verification
- 06Delivery
01
Brief
Tell us the use case, sector, language, sensitivity and how you intend to use the result. Two paragraphs is usually enough to start.
Step 102
Feasibility
We come back with what's possible, what isn't, and what an honest dossier for this dataset would have to disclose. If it's the wrong tool for the job, we'll say so.
Step 203
Quote
Scope, timeline and cost in writing. Includes the licence posture we can commit to and any reviewer notes that will attach to the dossier.
Step 304
Generation
Built by a verified provider on a method documented up front. You see samples at agreed checkpoints, not just at the end.
Step 405
Verification
Same review standard as catalog datasets — technical, legal and ethical — applied to your set. Anything that requires remediation comes back to the provider before delivery.
Step 506
Delivery
Dataset, dossier and executed licence delivered via secure link bound to your request. Versioning and audit trail recorded from day one.
Step 6Honest scope
where custom datasets earn their keep.
We'd rather decline a brief than over-promise. Here's where custom work is worth it — and where it isn't.
Good fits
- A scenario the public catalog doesn't cover (niche sector, rare language, specific edge cases).
- An existing dataset that almost fits but needs tuned coverage or a stricter licence posture.
- Evaluation sets purpose-built for your model's known failure modes.
Poor fits
- Anything that depends on real identifiable individuals — synthetic data is for testing systems, not adjudicating real cases.
- One-off, low-volume requests where the dossier overhead exceeds the value of the data.
- Use cases that need real production data signals; synthetic should validate, not replace.
Send a brief
two paragraphs is usually enough to start.
You don't need a finished spec. Tell us the use case and the rough shape; we'll come back with the questions that matter.
Built for teams who own the outcome
see what a documented dataset looks like before you brief.
Open a real listing and read its dossier — same structure your custom dataset will follow.