CVAT icon

CVAT

認領

CVAT is a data annotation platform for computer vision teams that work with images, video, and 3D data. It supports manual and AI-assisted labeling, quality review, collaboration, and dataset export for model training.

CVAT

What CVAT does

CVAT is a data annotation platform for computer vision teams working with images, video, and 3D data. It is designed to turn raw visual data into model-ready training datasets using a mix of manual labeling, AI-assisted annotation, review workflows, and dataset export tools.

The platform supports the full annotation lifecycle: importing data from local files or cloud storage, organizing work into projects and tasks, assigning contributors, validating results, and exporting finished datasets in formats used for model training. The product is available as CVAT Online, has an enterprise self-hosted offering, and is supported by a community version and open-source ecosystem.

Core capabilities

Flexible data import

Bring in local files, existing annotations, and cloud-stored datasets from Amazon S3, Azure Blob Storage, Google Cloud Storage, S3-compatible buckets, or self-hosted storage.

Project and team management

Create and manage projects, tasks, jobs, roles, permissions, stages, statuses, and ownership to keep annotation work organized at scale.

Manual and automated labeling

Label with boxes, polygons, masks, skeletons, tags, and tracks, then use AI-assisted tools such as SAM 2, SAM 3, Ultralytics, Hugging Face models, or custom models to speed up work.

Quality control workflow

Review jobs through validation and acceptance stages and use ground truth, honeypots, and consensus workflows to trace quality decisions.

Monitoring and analytics

Track progress, workload, and completion status across projects, tasks, and jobs to identify bottlenecks and measure throughput.

Export and workflow automation

Export datasets in 20+ formats and connect workflows through API, SDK, and CLI access for automation and integration.

Common use cases

  • Computer vision model training

    Prepare training datasets for detection, segmentation, tracking, and pose estimation by combining manual labeling with AI-assisted tools and export-ready formats.

  • Team-based labeling operations

    Coordinate annotation work across multiple contributors by assigning tasks, defining permissions, and reviewing jobs through validation and acceptance stages.

  • Dataset preparation and handoff

    Import data from cloud storage, work through labeling and QA, and export finished datasets into formats such as COCO, YOLO, KITTI, Cityscapes, or Pascal VOC.

  • Enterprise-controlled deployment

    Use self-hosted deployment when the organization wants CVAT inside its own infrastructure rather than a hosted online workspace.

  • Workflow automation

    Automate recurring annotation workflows with API, SDK, CLI, and AI-tool integrations for internal machine-learning pipelines.

Pros and Cons

Pros

  • Supports a broad set of visual data types, including images, video, and 3D data.
  • Combines manual annotation with AI-assisted labeling options to speed up repetitive work.
  • Includes collaboration, permissions, validation, and task tracking for team workflows.
  • Offers multiple export formats and programmatic access through API, SDK, and CLI.
  • Provides both cloud-hosted and self-hosted enterprise options.

Cons

  • The pricing and support pages indicate that some capabilities vary by plan, so teams need to compare tiers before choosing a deployment.
  • The source pages do not provide a detailed public breakdown of every supported connector or workflow integration.

FAQ

How does CVAT fit into an annotation workflow?

CVAT provides a browser-based workspace for importing data, organizing projects and tasks, labeling with manual and automated tools, validating annotations, and exporting datasets for model training. The source pages do not describe a separate installation wizard or onboarding flow for the online product.

What kinds of data and outputs does CVAT support?

CVAT supports images, video, and 3D data, along with common computer-vision tasks such as detection, segmentation, tracking, and pose estimation. It also supports exporting datasets in 20+ formats, including COCO, YOLO, KITTI, Cityscapes, and Pascal VOC.

Can CVAT be used by individuals and teams?

The platform is built for teams as well as solo users. The pricing page distinguishes Solo and Team plans, and the company site describes project, task, role, and permission management for collaborative work.

What deployment options are available?

CVAT Online is available through tiered plans, including a free plan and paid Solo and Team plans. The enterprise offering is self-hosted and positioned for organizations that want to run CVAT within their own infrastructure.

What integrations and support channels are available?

The sources show CVAT supports cloud storage connections, APIs, SDKs, CLI access, AI-assisted annotation, and integrations with Hugging Face and Roboflow. The support page directs product questions to Docs, GitHub, Discord, or the help desk depending on the product version.

Quick Facts

Category
Data annotation platform
Primary users
Computer vision teams, solo labelers, and enterprises
Data types
Images, video, and 3D data
Deployment
CVAT Online, community/open-source, and self-hosted enterprise
Integrations
Cloud storage, Hugging Face, Roboflow, API, SDK, and CLI
Source domain
cvat.ai