annotation tools
CVAT vs Labelbox vs V7: Computer-Vision Tools (2026)
CVAT vs Labelbox vs V7 for computer-vision annotation: open-source vs enterprise, auto-labeling, DICOM, and pricing models, compared. See which fits.

CVAT vs Labelbox vs V7 is the choice most computer-vision teams weigh once they outgrow a spreadsheet of labels, and the three occupy genuinely different positions: CVAT is the open-source tool you self-host, Labelbox is the enterprise data engine for large multimodal and RLHF programs, and V7 is the premium specialist for medical imaging and documents. The decision rarely comes down to who can draw a bounding box—all three do that well. The real deciders are four: open-source control versus managed scale, how each platform auto-labels, whether you need DICOM and compliance, and how you pay. This guide lines all three up fairly, with facts verified as of July 2026.
CVAT vs Labelbox vs V7: which platform is best?
There is no single best computer-vision annotation platform; the right one depends on whether you need open-source control, managed scale, or clinical and document specialization. CVAT wins when you want to own the stack and label image, video, or 3D data without a per-seat fee. Labelbox wins when you are running a large, data-centric program and want foundation-model auto-labeling plus a managed workforce. V7 wins when your data is medical or document-heavy and compliance is non-negotiable.
The through-line is the open-source-versus-commercial trade-off applied to vision. CVAT gives you control and zero license cost but hands you the infrastructure burden; Labelbox and V7 absorb that burden in exchange for subscription or usage fees and some lock-in. That framing decides most of what follows, and it is the subject of our deeper piece on open-source vs commercial annotation platforms.
CVAT vs Labelbox vs V7: the comparison table
Here is the artifact most CV teams want: the three platforms on the axes that decide the purchase. Read across a row for one platform; read down a column to compare all three on one dimension.
| Axis | CVAT | Labelbox | V7 (Darwin/Go) |
|---|---|---|---|
| License | Open source (MIT); paid cloud/Enterprise | Proprietary | Proprietary |
| Hosting | Self-host, managed cloud, or Enterprise | Cloud (air-gapped option; no classic on-prem) | Cloud (SaaS) |
| Modalities | Image, video, 3D/point-cloud | Multimodal: image, video, text, geo, DICOM, LLM | Image, video, DICOM, whole-slide, PDF/docs |
| Auto-labeling | SAM/SAM2, HF & Roboflow models, AI Agents | Model Foundry (foundation models); Catalog curation | Auto-Annotate (SAM2, MedSAM); video auto-track |
| Managed workforce | Labeling services available | Alignerr / Alignerr Connect (expert workforce) | Available via enterprise engagements |
| Compliance | Your responsibility when self-hosted | SOC 2 Type II, HIPAA, ISO 27001, GDPR | SOC 2 Type II, HIPAA, ISO 27001, GDPR |
| Pricing model | Free OSS · per-user cloud · Enterprise quote | Usage-based (Labelbox Unit ~$0.10); free tier | Enterprise quote; no public pricing / free tier |
| Best for | Teams wanting OSS control over image/video/3D | Large multimodal & RLHF data programs | Medical imaging & document-AI enterprises |
For the broader field beyond these three—including text-first and general tools—see our roundup of the best data annotation tools; for text-first tooling specifically, compare Label Studio vs Prodigy vs doccano.
Is CVAT free and open source?
CVAT’s Community edition is free and open source under the MIT License, so you can self-host it and use it commercially at no license cost. It began inside Intel in 2017, was maintained under the OpenCV umbrella, and spun out as the company CVAT.ai in 2022, which now offers a hosted CVAT Online edition and a supported Enterprise edition alongside the free community build.
The honest catch is that “free” is a license fact, not a total-cost fact. Self-hosting CVAT means running a Docker-based stack, scaling it for concurrent annotators, and owning upgrades and backups; the advanced analytics, audit logs, and governance features sit in the paid tiers. For a team without infrastructure staff, the managed cloud can be cheaper once you count engineering time—the classic open-source total-cost-of-ownership question covered in the open-source vs commercial comparison.
How do they auto-label?
All three lean on foundation models to pre-label, but they package it differently. CVAT integrates the Segment Anything Model (SAM and SAM2) for click-to-segment masks, plus Hugging Face and Roboflow model integrations and custom models via CVAT AI Agents, with batch auto-annotation on higher tiers. It is powerful and open, but you often wire up and manage the models yourself.
Labelbox centers auto-labeling in Model Foundry, which runs pre-trained and foundation models to pre-label data, alongside Catalog for embeddings-based search and curation and a Model product for evaluation and error analysis—a full data-centric loop of curate, label, evaluate, improve. V7’s Auto-Annotate is built on SAM2, with MedSAM for medical imagery and automatic object tracking across video frames; V7 states this can accelerate labeling substantially, which is a vendor claim rather than an independent benchmark. Across all three, the reliable pattern is the same: the model drafts, and a human verifies each mask or box.
Which is best for medical imaging and documents?
V7 is the strongest of the three for medical imaging and document-heavy work. It supports DICOM and whole-slide pathology images, offers MedSAM-based auto-annotation, and carries deep compliance—SOC 2 Type II, HIPAA, ISO 27001, and GDPR—the combination regulated healthcare and life-sciences teams need. Its separate V7 Go product extends into GenAI document workflows: applying foundation models to extract fields from printed and handwritten documents, with conditional routing of sensitive cases to human review.
Labelbox also handles DICOM and documents within its broader multimodal platform, and carries deep enterprise compliance (SOC 2 Type II, HIPAA, ISO 27001, GDPR), so it is a credible medical-imaging choice at scale, especially if you also need text and LLM workflows. CVAT is built around general image, video, and 3D annotation rather than clinical formats, so for DICOM-native pipelines it is the weakest fit of the three. If medical imaging is your whole use case, the specialization gap is real and worth paying for.
Pricing models compared
The three price on completely different models, so compare the model, not a number. CVAT is free to self-host; its managed CVAT Online is a per-user subscription and Enterprise is a quote. Labelbox uses usage-based pricing built on the Labelbox Unit at roughly $0.10 per unit, with different conversion rates per product (labeling consumes far more units than curation), a free tier for small teams, a pay-per-unit self-serve subscription that unlocks SSO, and enterprise quotes for HIPAA and volume discounts.
V7 publishes no pricing and has no free tier; it is enterprise sales-led and usage-based, so you go through a quote. Two cautions follow. First, Labelbox’s managed human-labeling services (Alignerr and Alignerr Connect) are quoted separately and can dwarf the platform fee, so the LBU number is only part of the bill. Second, usage-based and quote-based models are genuinely hard to forecast; model your real annotation volume before committing, and treat any specific figure here as a starting estimate to verify.
Where does Tagaroo fit?
Tagaroo is not a computer-vision platform and does not compete with CVAT, Labelbox, or V7 for bounding boxes, segmentation masks, or DICOM pipelines. If your job is labeling pixels at scale, one of those three is the right tool, and this section is the honest boundary of where we are useful.
Tagaroo works on language, not images: guided, template-driven annotation of clinical and research transcripts against a defined coding scheme, with inter-rater reliability built in. Where V7 labels a chest X-ray, Tagaroo codes the interview about the patient—tagging depressive severity against the MADRS or disorganized speech against the TLC scale, with every rating anchored to the exact span that justifies it. It is a different job entirely. If your data is visual, use a vision platform; if it is spoken or written clinical language coded against an instrument, that is Tagaroo’s lane.
Which should you choose?
Pick the platform for your data and your constraints, not its market share. The three map onto three clear situations.
- Choose CVAT if you want open-source control over image, video, or 3D annotation, have (or can spare) infrastructure staff, and prefer no per-seat fee. Accept that self-hosting, scaling, and governance are on you, and that the polished analytics live in paid tiers.
- Choose Labelbox if you are running a large, data-centric or RLHF program that needs foundation-model auto-labeling, curation, mature compliance, and an on-tap expert workforce. Accept complex usage-based pricing and no traditional on-prem.
- Choose V7 if your data is medical (DICOM, whole-slide) or document-heavy, you need HIPAA and polished auto-annotation, or you want GenAI document automation via V7 Go. Accept premium, quote-only pricing and a smaller community than the open-source option.
The practical upshot of the CVAT vs Labelbox vs V7 decision: match the platform to your modality, your compliance needs, and your appetite for running infrastructure. For pixels, that is one of these three; for coding clinical language against a validated scheme with reliability from the first label, start a scale-based project in Tagaroo instead.
References
- CVAT.ai. CVAT repository and license (MIT). GitHub · Pricing · Docs overview
- OpenCV. CVAT has a new home (Intel origin, OpenCV, CVAT.ai spin-out), 2022. opencv.org/blog
- Labelbox. Plans and pricing (Labelbox Unit model). labelbox.com/pricing
- Labelbox. Privacy and security program (SOC 2 Type II, HIPAA, ISO 27001, GDPR; air-gap). labelbox.com/company/security
- V7 Labs. V7 Darwin (modalities, SAM2/MedSAM, auto-track, compliance). v7labs.com/darwin
- V7 Labs. DICOM and medical imaging annotation. v7labs.com/dicom-annotation-tools
- V7 Labs. Pricing and V7 Go. v7labs.com/pricing
If your annotation is clinical language rather than images, Tagaroo turns instruments like the MADRS and the TLC scale into guided, evidence-anchored coding workflows with inter-rater reliability computed as your coders work.
Frequently asked questions
- Is CVAT free and open source?
- Yes. CVAT's Community edition is open source under the permissive MIT License, so you can self-host and use it commercially for free (CVAT.ai, 2026). CVAT.ai also sells a managed cloud edition (CVAT Online) and a supported self-hosted Enterprise edition. Labelbox and V7 are both proprietary, closed-source platforms with no free self-hosting option.
- What is the difference between CVAT, Labelbox, and V7?
- CVAT is the open-source, self-hostable computer-vision tool, strong on image and video for teams that want control and no per-seat fee. Labelbox is a commercial, cloud, data-centric platform with foundation-model auto-labeling and a managed expert workforce, aimed at large multimodal and RLHF programs. V7 is a premium commercial platform specialized in medical imaging (DICOM), documents, and GenAI document workflows via V7 Go.
- Which annotation tool is best for medical imaging and DICOM?
- V7 is the strongest of the three for medical imaging: it supports DICOM and whole-slide images, offers MedSAM-based auto-annotation, and holds SOC 2 Type II, HIPAA, ISO 27001, and GDPR compliance (V7 Labs, 2026). Labelbox also handles DICOM within its multimodal platform. CVAT focuses on general image, video, and 3D annotation rather than clinical imaging formats.
- How much do CVAT, Labelbox, and V7 cost?
- CVAT is free to self-host (MIT); its managed cloud is a per-user subscription and Enterprise is a quote. Labelbox uses usage-based pricing built on the Labelbox Unit (about $0.10 per unit), with a free tier, a pay-per-unit subscription, and enterprise quotes. V7 publishes no pricing and has no free tier—it is enterprise sales-led and usage-based. Exact figures change often, so confirm current terms before buying.
- Does CVAT do auto-labeling?
- Yes. CVAT integrates the Segment Anything Model (SAM/SAM2) for click-based segmentation, plus Hugging Face and Roboflow models and custom models through CVAT AI Agents, with batch auto-annotation on higher tiers (CVAT.ai, 2026). Labelbox auto-labels via Model Foundry (running foundation models), and V7 via Auto-Annotate (SAM2, and MedSAM for medical imagery).
Put this into practice
Tagaroo turns any rating scale or coding scheme into a guided annotation workflow — with inter-rater reliability computed as you go.