Ultrasound Data Collection and Annotation for AI: Why This Modality Demands a Different Approach
A look at the unique data collection and annotation challenges facing AI developers who build with ultrasound – and what a proactive approach actually looks like.

The Modality No One Is Storing
For AI developers the hardest medical imaging data to acquire is ultrasound.
It is not rare – ultrasound is one of the most widely performed diagnostic procedures. But the problem is what happens to the data afterward.
With CT and MRI, images are archived in PACS systems as a matter of course. The DICOM files sit on hospital servers, retrievable for research, retrospective studies, and AI training. But with ultrasound the standard workflow is different: the sonographer performs the scan, captures a handful of representative screenshots, and hands the patient a printout or a PDF report. The dynamic exam – the actual imaging session, the probe movements, the structures visualized in real time – is rarely stored in any structured database.
The practical consequence for AI developers is that you cannot build an ultrasound AI model from passively accumulated hospital data the way you can with CT or MRI. To get the data we have to be proactive.
Why Ultrasound Is Still the Frontier for Medical AI
The global Ultrasound AI Market is valued at $2.09 billion in 2025 and is expected to reach $22.03 billion by 2035, growing at a CAGR of 26.62% (according to Wikipedia). The demand is growing: healthcare systems face various challenges from increasing patient volumes, escalating operational costs, and a scarcity of skilled sonographers. As a result – extended patient wait times and professional burnout.
By 2035, key applications are expected to span cardiology, obstetrics, orthopedics, and gastroenterology. AI-powered solutions become pivotal in early disease detection, workflow optimization, and reducing reliance on highly specialized radiologists.
The Three Pillars of Ultrasound AI Data

Thyroid Imaging: High-Stakes Decisions, Standardized Labels
Thyroid ultrasound is the most mature category for AI model development. The TI-RADS (Thyroid Imaging Reporting and Data Systems) classification provides sonographic criteria for describing the suspicious findings of thyroid nodules, developed to characterize the risk of malignancy and better guide decisions around fine-needle aspiration cytology.
The ACR TI-RADS system (the version most commonly used in AI development today) evaluates five key sonographic features of a nodule: composition, echogenicity, shape, margin, and echogenic foci.
Each feature carries a specific point value. The points are summed to determine a TR level from TR1 (benign, no FNA needed) through TR5 (highly suspicious, FNA if ≥1 cm). The scoring system is detailed and reproducible. Also,it provides a structured annotation framework that links visual features directly to clinical risk which is critical for AI.
Deep learning-based methods can provide powerful assistance to radiologists, but their performance depends on the quantity and quality of training data. Most existing thyroid ultrasound datasets either rely on TI-RADS assessments alone as labels or are simply not publicly available. This creates a significant bottleneck for developers.
We have a dual challenge: high-quality thyroid datasets need not just images, but images annotated by radiologists who understand TI-RADS scoring in clinical practice. High quality deep learning frameworks fine-tuned for automated nodule detection, segmentation, and malignancy classification aligned with ACR TI-RADS criteria require expert-annotated regions of interest to achieve high level of sensitivity and specificity. The difference in annotation quality matters enormously for model performance. We have discovered that ourselves while doing thyroid ultrasound annotation for a US-based healthcare AI company.
Pregnancy and Fetal Pathology: Sensitive Data, Proactive Collection
Obstetric ultrasound is one of the highest-volume imaging procedures globally, performed on virtually every pregnant patient multiple times throughout a pregnancy. The International Society of Ultrasound in Obstetrics and Gynecology recommends routine obstetric ultrasounds between 18 and 24 weeks’ gestational age to confirm pregnancy dating, measure the fetus so growth abnormalities can be recognized quickly, and assess for congenital malformations.
The challenge with fetal ultrasound data for AI training is that the most clinically valuable cases (involving fetal pathologies, structural anomalies, or early markers of developmental disorders) are precisely the ones least likely to be systematically collected. Standard prenatal care data is abundant but often lacks the annotated pathology cases an AI model needs for training.
This is where a proactive, clinic-partnership model becomes essential. Working directly with specialty clinics that see a high concentration of fetal pathologies makes it possible to build datasets that reflect the diagnostic edge cases an AI tool needs to be useful (especially if these facilities handle complex cases or high-risk pregnancies). These cases are not waiting in a database. They have to be identified, consented, collected, and labeled by clinicians who understand both the imaging findings and the clinical context.
What makes this category particularly sensitive. On top of the logistical challenges, pregnancy ultrasound data carries a distinct privacy weight that sets it apart from most other imaging modalities. These images can contain a significant amount of personal information: the name of the pregnant person, the hospital, the date and time of the procedure, gestational age, and the sonographer’s name. They can also reveal potential malformations or genetic conditions – information that is profoundly sensitive and potentially consequential for both the patient and the child. Obtaining proper, informed consent is therefore not just a regulatory formality; it is a genuine ethical obligation. That’s why in many cases the proactive collection with a trusted clinical partner is the only viable path.
On the question of fetal faces. One question that occasionally comes up in practice is whether a fetal face visible in an ultrasound image makes that image identifiable – and whether it should therefore be blurred as part of anonymization. This is a nuanced area. Standard 2D clinical ultrasound images of fetal faces do not carry the same re-identification risk as adult portrait photographs. The acoustic rendering, image resolution, and lack of cross-referenceable features make facial recognition from routine 2D scans practically implausible with current technology. The primary identifiers in a pregnancy ultrasound are the overlay text (patient name, date, hospital, gestational age) and DICOM metadata – and those are the elements where robust de-identification must be applied first.
The situation is different for 3D and 4D ultrasound imaging, where surface-rendered fetal faces are considerably more realistic and detailed. These images are more widely shared on social media as keepsake content, which increases the possibility of cross-referencing with other personal data. In a research or AI training context, blurring 3D/4D fetal faces is a reasonable and defensible precaution, even if formal guidance remains limited. For standard 2D clinical imaging used in AI development, the emphasis should be on thorough metadata removal and overlay text redaction, with fetal face blurring remaining an optional additional layer for teams that want maximum caution. As regulatory frameworks around biometric data continue to evolve, erring on the side of more thorough de-identification is always the safer position.
Women’s Health and Mammological Ultrasound
Breast ultrasound occupies an important complementary role to mammography, particularly for women with dense breast tissue where mammography sensitivity is reduced. It can be performed for either diagnostic or screening purposes, and is especially useful for younger women with denser fibrous breast tissue that makes mammograms more challenging to interpret.
The BI-RADS (Breast Imaging-Reporting and Data System), published by the American College of Radiology, standardizes reporting and helps clinicians communicate a patient’s risk of developing breast cancer – particularly relevant for patients with dense breast tissue. BI-RADS was actually the template on which TI-RADS was later modeled, reflecting the same philosophy of structured, reproducible risk communication that makes annotation for AI training tractable.
For AI developers in the women’s health space, the annotation challenge is similar to thyroid: you need radiologists or sonographers who work with BI-RADS in practice, not just in theory, to produce labels that a model can actually learn from.

What Ultrasound Annotation Actually Involves
Ultrasound annotation is meaningfully different from annotating CT or MRI data, and here is why.
The images are noisier. Speckle artifact, shadowing from calcifications, and acoustic enhancement behind cysts are all features that an experienced sonographer reads as information. A novice annotator might either miss or misattribute that. The same structural finding can look quite different depending on probe angle, frequency, and gain settings.
This is why annotation quality depends heavily on the clinical experience of the annotators. Labeling a thyroid nodule according to ACR TI-RADS requires understanding the five compositional features and applying them correctly under real-world image variation. Marking “punctate echogenic foci” (which score 3 points and push a nodule toward TR4 or TR5) requires knowing the difference between true foci and acoustic artifacts that can mimic them.
For body structure annotation, the challenge shifts slightly: it’s less about risk stratification and more about precise delineation of anatomical boundaries under image conditions that make those boundaries ambiguous. Annotation teams that work with ultrasound daily, rather than occasionally, develop calibration that takes time to build and cannot be easily substituted with general medical image annotation experience.
AI models for thyroid ultrasound that provide human-understandable features: texture, margin, echogenicity, shape, and location. They are more clinically accepted and significantly more interpretable than models that produce only a single risk score. Achieving this requires training data with feature-level annotations, not just binary labels. The annotation work is correspondingly more demanding.

Why Getting This Data Is Harder Than It Sounds
Even with the right clinical relationships and a sound collection protocol, building a useful ultrasound dataset involves friction at almost every step. Often they are invisible until you’re already in the process.
Hospitals and clinics that do have structured ultrasound databases often have them as a byproduct of specific research programs or an unusually rigorous data governance culture, not as a standard institutional practice. Identifying which facilities have usable data, negotiating access, aligning on de-identification requirements, and establishing data transfer protocols all take time that is easy to underestimate.
Consent workflows add another layer. For pregnancy data in particular, consent must cover both the patient and, in many jurisdictions, the potential future interests of the child – a consideration that few standard consent forms are designed to address. Getting consent right requires working with clinical and legal teams at each partner institution, not just applying a template.
Then there is the image quality problem. Ultrasound image quality varies significantly across devices, operators, and settings. Retrospective data collected from clinical practice often includes images taken at sub-optimal angles, with poor gain settings, or from equipment with older transducer technology. For AI training, this variability is both a challenge and an advantage at the same time. Models trained on heterogeneous real-world data generalize better but it means that quality control cannot be automated away. Every dataset requires human review at a scale that most teams don’t anticipate going in.
Finally, annotation bottlenecks are real. A clinical expert who can annotate thyroid nodules to TI-RADS standards is not interchangeable with a general radiologist, and they have full clinical schedules. Building annotated datasets at meaningful scale requires either long-term relationships with specialist annotators or an in-house team with the right clinical background. But even then, throughput is limited by the complexity of the task.
The Retrospective/Proactive Balance
Any ultrasound data strategy has to grapple with two collection modes, and the right balance depends entirely on the project.
Retrospective collection works when a clinic has already built a database – whether because of an atypically rigorous data management culture, a research program, or a specialty focus that created documentation habits above the norm. These databases exist, but finding them requires the right clinical relationships. The data is there; it just isn’t waiting on a public server.
Proactive collection is necessary when the retrospective pool is insufficient in volume, pathology distribution, or annotation quality. This means designing a collection protocol, establishing consent procedures that comply with GDPR and HIPAA requirements, and embedding collection into clinical workflows in a way that doesn’t compromise patient care. Done well, it produces exactly the data a model needs. Done poorly, it produces data that looks right on the surface but has systematic biases that undermine model performance.
The critical insight is that for most serious ultrasound AI development, some degree of proactive collection is unavoidable. The data simply doesn’t exist in a form ready for training without deliberate work to create it.

Why This Matters Now
Research has consistently identified inadequate implementation infrastructure, limited workflow integration, and the absence of ongoing performance monitoring as primary barriers preventing clinically validated AI tools from reaching routine use beyond their originating institutions.
For ultrasound specifically, a fourth barrier often precedes all the others: training data that is thin, poorly labeled, or clinically unrepresentative. Models trained on publicly available datasets often fail when deployed in real hospital environments where image quality, case mix, and workflow context are different.
The ultrasound AI developers who are advancing most reliably are the ones who treat data collection and annotation as a core competency rather than a procurement problem.
medDARE supports ultrasound data collection and annotation for AI development programs. Our team works with partner clinics for both retrospective access and proactive collection, and our annotation specialists have 5+ years of hands-on experience labeling ultrasound pathologies and anatomical structures.
→ Related reading: AI in Women’s Ultrasound: The Power of Medical Data






















