NeurIPS 2026 Evaluations & Datasets

CoMET-Bench + CoMET-Agent

Find every event.
Keep the evidence.

Conditional Multi-Event Temporal Grounding in Long-Form Video

Search a long video for every event that meets your conditions. Return the count—and the timestamps that support it.

Yuanhao Zou and collaboratorsAll authors ↓

University of Central Florida × Qualcomm AI Research

A question becomes evidencePaper example

Between 16:00 and 34:00, how many distinct segments show a jaguar swimming against the river current?

HF · CoMET-Bench

Video files require approved HF access. Request access ↗

Search within the condition16:00–34:00
00:0044:00
Verified momentsDetail view
123 30:4031:50
Answer

3distinct events

Ground-truth examples from the paper and HF dataset; no live model inference.

600long videos
2,789conditional queries
33.8 minaverage video length
5real-world domains

CoMET-Bench

One query.
Every matching moment.

Conditions can span time, space, and identity. The benchmark measures whether a model can find all qualifying events, count them, and recognize when none exist.

SportsTV & movieLife recordKnowledgeSurveillance

Count

How many distinct events meet the query?

Ground

Where does each event start and end?

Reject

Rejection-F1 rewards correct empty answers while penalizing always-empty predictions.

Video sources & annotation quality
11annotators791person-hours (approx.)93.5%direct consensus at tIoU > 0.7
SourceVideosShare
Video-MME21435.7%
MLVU18030.0%
CG-Bench19833.0%
YouTube81.3%

All 2,789 queries were independently double-annotated; 182 queries required adjudication. The final benchmark contains 15,022 intervals.

CoMET-Bench · real video examples

See the condition. Find the moments.

All example videos on HF ↗

These panels match the queries and ground-truth intervals shown in Figures 1–2. Select a condition, then click a timestamp to watch the corresponding event.

Video files require approved HF access. Request access ↗

Jaguar swimming

Figure 1(d) / Figure 4
3annotated events
BoundedDynamic

Between 00:16:00 and 00:34:00, how many distinct segments occur where the jaguar is shown performing a swimming motion against the river current?

Click a timestamp to play its annotated segment.
5WIdIs3A9Ok_q001Video on HF ↗

Scoring through a foul

Figure 2(a)
5annotated events
DynamicSynchronous

Count the number of times a player successfully scores a field goal while a foul is simultaneously committed by the defense, resulting in an 'and-one' play.

Click a timestamp to play its annotated segment.
KLaBd-0yCE4_q001Video on HF ↗

Bench press after 120 lb

Figure 2(c)
5annotated events
DynamicStaticSequential

After the weight reaches 120 LBS, find all the moments that a person wearing sunglasses successfully complete the bench press challenge.

Click a timestamp to play its annotated segment.
xre_3dxdhkQ_q004Video on HF ↗

A cut back to the speaker

Figure 2(d)
7annotated events
BoundedDynamicIdentity

Between 00:00:00 and 00:10:00, count the number of times the video cuts from a clip of the TV show back to the speaker in the blue suit sitting in front of the camera.

Click a timestamp to play its annotated segment.
mzvHYBgD-tI_q003Video on HF ↗

News under a causal condition

Figure 2(b)
2annotated events
StaticCausal

Identify and ground the moments in the video where the reported events are influenced by the U.S.–Iran conflict.

Click a timestamp to play its annotated segment.
7NNAw_zvvOc_q002Video on HF ↗

A plausible event that never occurs

Figure 2(e)
0annotated events
IdentityNegativeBoundedDynamic

Between 00:00:00 and 00:05:00, how many times does a male adult enter the room to take care of the baby?

No qualifying interval: [ ]
Context preview inside the first five minutes; the annotated answer is empty.
0175_surveil_23_q003Video on HF ↗

CoMET-Agent

Turn the video into a search.

Training-free

The Planner Agent interprets the query, selects graph settings, and narrows the search to any explicitly bounded time window.

CoMET-Agent architecture: planning, hierarchical graph construction, iterative verification, and final aggregation from global memory.
CoMET-Agent architecture · Figure 4 in the paperView full size ↗

Grounding results

Same backbone.
Better grounding.

Structured search improves over single-pass inference across three evaluated backbones. Select a model to compare the reported results.

CoMET-Bench · positive queries · F1@tIoU 0.5 (%) · main-results table

What remains challenging?

Fine-grained entity tracking, uniform retrieval across long videos, and pairing causes with their consequences remain open challenges.

Grounding F1@0.5

+6.1percentage points

GPT-5 · single pass10.1%
CoMET-Agent · GPT-516.2%
081624%

Reported benchmark values, not live measurements.

Complete evaluation of additional agents

n = 2,789

Completed runs supersede the earlier partial results. F1 is evaluated on positive queries; rejection metrics include the negative-query subset. Adapted methods retain their search pipeline and change terminal output formatting.

MethodBackboneF1@0.5 ↑Rej.-F1 ↑FPR ↓
T*-adaptedGPT-4o1.434.30.0
Vgent-adaptedQwen2.5-VL2.048.55.4
TimeSearch-RQwen2.5-VL5.026.183.7
CoMET-AgentQwen3-VL8.857.48.0
Reported per-query latency
MethodBackboneSeconds / query
CoMET-AgentQwen3-VL21.9
CoMET-AgentQwen2.5-VL27.5
TimeSearch-RQwen2.5-VL49.1
Vgent-adaptedQwen2.5-VL87.0
CoMET-AgentGPT-4o21.3
T*-adaptedGPT-4o207.8
VideoARMGPT-4.1 / o3287.2

Reported wall-clock times follow the hardware and API setups in the paper; these are not hardware-normalized algorithmic speedups. CoMET-Agent uses DINOv2 ViT-L/14.

The team behind CoMET

Yuanhao Zou1Arthad Kulkarni1Lucas Toñanez1Lincoln Spencer1Guangyu Sun1Tianxingjian Ding1Andong Deng1Yi Li2Shuangjun Liu2Yuan Li2Dashan Gao2Ning Bi2Taotao Jing2Shuai Zhang2Chen Chen1,†
1 University of Central Florida2 Qualcomm AI Research† Corresponding author
Acknowledgments & funding

This work was supported in part by a gift from Qualcomm. This work also used the Delta system at the National Center for Supercomputing Applications, which is supported by the National Science Foundation under Award No. OAC-2005572, through allocations CIS260306 and CIS250828 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program. ACCESS is supported by the U.S. National Science Foundation under Grants No. 2138259, 2138286, 2138307, 2137603, and 2138296.

The authors thank Francisco Arisso, Ayaan Khan, David Orjuela, Harshitha Sathees Kumar, and Jeremy Triana for their contributions to the annotation of the data.