The Real Problems AI Companies Run Into When Annotating Digital Pathology Data

The Real Problems AI Companies Run Into When Annotating Digital Pathology Data
title
title

Accelerating your AI Success

Explore
July 22, 2026 | 7 min read

 

The Real Problems AI Companies Run Into When Annotating Digital Pathology Data

 

Digital Pathology Annotation

Most teams building AI for digital pathology underestimate the annotation stage. It gets scoped like any other computer-vision labeling job – draw bounding boxes or outlines, assign classes, ship the dataset. Then the project actually starts, and it turns out pathology annotation breaks in ways that generic image labeling doesn’t. The problems below are the ones that come up over and over, usually after a team has already committed to a timeline and a budget that didn’t account for them.

 

Problem 1: Picking the wrong level of annotation for what the model actually needs

There are two fundamentally different types of annotation in digital pathology, and choosing between them isn’t a style preference – it’s a decision that determines whether the resulting dataset can support the model at all.

Polygonal segmentation means outlining a whole region or cluster of cells and giving it one label: “this is tumor,” “this is stroma,” “this is necrotic tissue.” It’s fast and cheap, and it works well for tissue-level classification or region-of-interest detection.

Cell-level segmentation and classification is a different task entirely. A specialist marks individual cells – sometimes just the nucleus, sometimes nucleus, cytoplasm, and membrane as separate structures – and assigns each one to a class. Teams that need cell-level output (biomarker scoring, immune cell quantification, anything requiring per-cell precision) but scope the project as polygon work end up with a dataset that simply can’t answer the question their model was built to answer. That mismatch usually isn’t discovered until the model underperforms and someone traces it back to the training data.

Comparison of Common Medical Annotation Types (2)

 

Problem 2: Underestimating how granular the class taxonomy needs to be

Even teams that correctly scope for cell-level annotation often underestimate how many classes that actually requires. A PD-L1 scoring project, for example, can run to a dozen-plus classes:

  • Neoplastic cells, split by PD-L1 status (positive/negative)
  • Neoplastic membrane staining, split by completeness (complete circumferential vs. partial)
  • Lymphocytes, split by PD-L1 status
  • Plasma cells, split by PD-L1 status
  • Fibroblasts
  • Endothelial cells
  • Granulocytes (covering neutrophils, eosinophils, and rare basophils)
  • Macrophages, split by positive/negative staining
  • Normal epithelium
  • An “ignore” bucket for genuinely ambiguous cells

This isn’t a taxonomy for its own sake. A PD-L1 combined-positive-score model has to learn to distinguish tumor cells with complete circumferential membrane staining from ones with partial staining, and to separate both from the immune cell populations sitting right next to them. Collapse that taxonomy into fewer, broader classes and the dataset stops representing the clinical distinction the model needs to learn. Notably, the tissues where this kind of granular scoring matters most – brain, breast, and prostate – are also where the clinical stakes are highest: breast and prostate cancers are among the most commonly diagnosed cancers worldwide, per WHO’s breast cancer fact sheet, and brain tumors, while rarer, are disproportionately hard to treat. Getting the taxonomy wrong on these tissues has real downstream costs.

Digital Pathology Data Annotation

Problem 3: Sourcing annotators who are actually qualified to make the calls involved

A lot of digital pathology annotation gets routed to generalist data-labeling teams, and a lot of it shouldn’t be. The judgment calls involved – is this a fibroblast embedded in collagen or something else, does this membrane staining count as complete or partial, is this smudge an artifact or a cell that should be flagged “ignore” – require the kind of training a certified pathologist has and a generalist annotator doesn’t.

This is a hard line worth holding: pathology annotation should be done by pathologists who are currently certified in the country where they trained, not by annotators working from a reference guide. CAP’s guidance on validating digital pathology systems exists precisely because the industry learned, the hard way, that moving pathology into a digital environment doesn’t lower the bar for who’s qualified to interpret it – if anything it raises the bar for documentation and reproducibility. medDARE holds this line internally as well: annotators on pathology projects carry active certification in their country of practice, drawing on specialists licensed across both Ukraine and the European Union.

Problem 4: Finding pathologists who have the time, not just the credentials

Even once a team accepts that certified pathologists are non-negotiable, there’s a second, harder problem: there simply aren’t very many pathologists in the world, and the ones who exist are already carrying full clinical caseloads. Asking an overworked pathologist to spend hours annotating cells for someone else’s model is, understandably, a hard sell – which is why so many AI teams stall out at the sourcing stage even after they’ve correctly identified what kind of specialist they need.

Digital Pathology Data Annotation

One pattern worth noting: pathologists in Ukraine, in particular, have often shown more availability and interest in this kind of project work than the broader specialist community – not because their clinical caseload pressure is any lower, but because there’s a genuine appetite among that group for contributing to AI development. For teams that hit the sourcing wall, that’s turned out to be a meaningful way through it.

 

Problem 5: Outgrowing the annotation tool partway through the project

QuPath is the default starting point for almost every digital pathology annotation project, and for good reason – it’s free, purpose-built for pathology, and its brush, wand, and polygon tools cover most needs without requiring custom software. Manual annotation, without model-assisted pre-labeling, is still the standard for ground-truth work, since the entire point is producing training data that hasn’t already been shaped by another model’s biases.

The problem shows up later: most teams eventually want an annotation environment tailored to their exact class taxonomy, QC process, and data pipeline, and QuPath wasn’t built to be that. A recent breast and prostate tissue project illustrates the pattern well – two dedicated pathology annotators working through nucleus, cytoplasm, and membrane segmentation, with a QC loop built specifically to catch labeling drift before it became systemic. That kind of workflow tends to emerge only after a team has already outgrown a general-purpose tool, which means the transition is usually reactive rather than planned for. For more on why pathology data quality is such a specific and demanding problem in medical AI, there’s a deeper look at that here.

Digital Pathology Data Annotation

Problem 6: Treating annotation as the end of the pathologist’s job

The final problem is more of a missed opportunity than a hard blocker: many teams budget for labeling and stop there, when the pathologists doing the annotation work are often well positioned to help with what comes after. They can validate the resulting model against cases it hasn’t seen. Just as usefully, they can flag product issues that have nothing to do with pixel-level accuracy – whether a review interface is actually usable in a real pathology workflow, or whether the tool is missing an entire class of finding that any working pathologist would consider table stakes. That kind of feedback loop tends to matter more to real-world adoption than another percentage point of model accuracy, and it’s routinely left on the table simply because it wasn’t scoped into the project from the start.

None of these problems are exotic – they show up on almost every digital pathology annotation project, in roughly this order, as a team moves from scoping to sourcing to execution. The teams that plan for them upfront tend to end up with usable datasets on the first pass. The teams that don’t tend to find out the hard way, usually several months and one underperforming model later.

Oleksandr Chyrkov
Expert author Oleksandr Chyrkov MRT Quality Manager
You may also like:

Want to know how we can accelerate your AI success?

Get a quote