Two clinicians reviewing a brain scan on diagnostic monitors

Medical imaging datasets

CT and MRI studies, profiled one by one, qualified to your spec.

More than nine thousand de-identified imaging studies with radiologist reports, each read at header and pixel level and scored against acceptance criteria before it is offered. Built for diagnostic AI teams that need to know exactly what they are training on.

9,500+
Studies
8.5M+
DICOM images
8,700+
CT studies
900+
MRI studies
87%
With radiologist report
91%
Lossless compression

Rounded inventory figures as of October 2026. Counts move as studies are indexed and qualified; exact eligible counts are returned against your specification.

The catalogue

Four collections ready to qualify

Each collection has its own page with composition, acquisition profile and qualification status. Smaller subsets are listed below and available on request.

Also in the catalogue, on request

Extremity, neck, spine, pelvis (CT)
1,000+
Brain, abdomen, prostate (MRI)
300+
Digital radiography
~50

New studies arrive weekly from the source network. If the body region or modality you need is not listed, ask: we run feasibility against the live catalogue, not this page.

Composition

A South Asian patient population across all adult age bands

Most public imaging datasets come from North American and European academic centres. This catalogue is drawn from public hospitals and diagnostic centres in India, which gives diagnostic models a population and a scanner mix they rarely see in training.

Patient sex, whole catalogue

  • Male60%
  • Female40%

Age band, whole catalogue

  • 0–1710%
  • 18–3934%
  • 40–5932%
  • 60–7921%
  • 80+3%

Catalogue at a glance

Acquisition years
2023 – 2026
Source institutions
25+
Scanner fleet
Siemens-dominant, with Philips, GE and United Imaging
Report pairing
One matched radiologist report per study where present
Delivery format
Complete DICOM studies, de-identified, to your bucket

How the catalogue is built

Every study is read, not just listed

A folder name says 'CT abdomen'. The pixel data says whether it is lossy, whether the portal-venous series is 1.5 mm or 5 mm, and whether there is a patient name burned into the corner. We index at that level, continuously, for every study that arrives.

  • Header and pixel-level indexing of every series, including transfer syntax and compression history
  • Duplicate detection by content hash and clinical near-match, so a study is counted once
  • Report pairing by accession and identifier, with unmatched reports held rather than guessed
  • Burned-in annotation detection with automatic quarantine
  • Acceptance verdicts recomputed whenever a specification or the corpus changes

Recorded on every study profile

  • Modality and body region
  • Contrast phases present and phase pattern
  • Thinnest and thickest slice
  • kVp range and bits stored
  • Compression and compression history
  • Manufacturer and scanner model
  • Field strength and sequence inventory (MRI)
  • Patient sex and age band
  • Acquisition year
  • Craniocaudal coverage
  • Report present and matched
  • Burned-in annotation check
  • Duplicate and multi-study folder check
  • Current acceptance verdict and failure codes

Qualification & delivery

From your inclusion criteria to a verified batch

The same process has delivered more than two thousand qualified CT studies to a diagnostic AI developer since mid-2026, in batches of around fifty, each verified against a hashed manifest.

01

Specification

We turn your inclusion criteria into a machine-checkable requirement set: modality, anatomy, contrast phase, slice thickness, kVp, bit depth, compression, report pairing and demographic constraints.

02

Index & profile

Every study is read at header and pixel level and profiled on more than thirty attributes, including phase pattern, scanner model, coverage, compression history and report match.

03

Verdict per study

Each study receives an eligible, review or reject verdict with the exact criterion codes it failed. Studies whose phase cannot be read from metadata go to a radiologist, not a guess.

04

De-identify & deliver

Eligible studies are de-identified, re-verified, packaged with a hashed manifest and shipped to your bucket in batches. You load your own pass/fail list and we reconcile against it.

Consent & permitted use

What this data may be used for, stated up front

Every study in the catalogue was obtained under a signed institutional research agreement with the source imaging network, with a data processing agreement on file. The permitted-use scope travels with the data and is written into every delivery manifest.

Permitted purpose

Development, validation and testing of diagnostic AI. Each delivery is matched to a declared purpose before a single study is packaged.

Permitted jurisdictions

United States and India. Deliveries to other jurisdictions require a fresh agreement with the source, which we will tell you plainly if it is not available.

Not permitted

General-purpose or foundation-model pretraining, and onward resale or sublicensing to third parties. We do not sell around the scope of the source agreement.

Retention

Time-limited retention is part of the agreement. Deliveries carry a retention term and we issue destruction certificates when a term ends or a study is withdrawn.

De-identification

DICOM headers are scrubbed to a documented profile, study, series and instance UIDs are remapped through a sealed crosswalk, and studies with burned-in annotation are detected and quarantined rather than shipped.

Provenance

Each batch ships with a hashed manifest, the acceptance verdict for every study and the source-agreement reference, so your counsel can trace the chain of rights.

Need the de-identification methodology and source-agreement summary for a vendor review? Request the compliance pack.

Who it is for

Built for teams that have to defend their training data

Diagnostic AI developers

Detection, segmentation and triage models

You have a specification and a validation plan. We return exact eligible counts against it, deliver in batches to your bucket, and reconcile your pass/fail audit against our ledger so disputes are settled on evidence.

Radiology AI vendors scaling to new markets

Domain shift and external validation

South Asian patients, Siemens-heavy fleets and public-hospital protocols are under-represented in most training sets. A qualified external cohort tells you how your model behaves before a customer does.

Research groups and CROs

Documented cohorts with audit trails

Datasheets, verdict ledgers, hashed manifests and destruction certificates are standard, so the cohort you train on is the cohort you can describe in a submission.

FAQ

Frequently asked questions

Where does the imaging data come from?

From a network of hospitals and diagnostic centres in India, supplied under a signed institutional research agreement with a data processing agreement on file. More than twenty-five source institutions are represented. We name the agreement and its scope in every delivery; we do not name the institutions publicly.

Is the data de-identified?

Yes. DICOM headers are scrubbed to a documented profile, every study, series and instance UID is remapped through a sealed crosswalk, and studies with burned-in text in the pixel data are detected and quarantined. Age is carried as a band. Reports are de-identified as text. Your compliance team receives the methodology document.

Can we license this for foundation-model pretraining?

No. The source agreement permits development, validation and testing of diagnostic AI in the United States and India. It does not permit general-purpose pretraining or onward resale, and we will not structure a deal that pretends otherwise. If you need data for pretraining we will say so and discuss sourcing it on a different basis.

How current are the figures on these pages?

They are rounded aggregates from the live catalogue as of October 2026. New studies are indexed continuously and figures move. When you send a specification we return exact eligible counts, not these rounded ones.

What do we actually receive?

Complete de-identified DICOM studies, every series, delivered in batches to a cloud bucket of your choice, with a hashed manifest per batch, the acceptance verdict for every study, de-identified report text where present and a datasheet describing composition and known gaps.

How quickly can we see a sample?

A sample batch against your specification is typically available within two weeks of the scoping call, after your compliance review of the methodology. Production cadence is agreed per engagement; the current client receives roughly fifty studies per batch, several batches per week.

Tell us what your model needs to see.

Send a specification, or just the clinical question. We will come back with the exact eligible count from the live catalogue, what the remainder fails on, and a sample batch. Figures on this page are as of October 2026.

Or email [email protected] · Mohali, India