AGENTIC VIDEO UNDERSTANDING

Understand video.Not just watch it.

Witness benchmarks AI agents that work out what happens in a video: the events, the dialogue, the cuts, the text on screen, the sounds.

The agent decides which frames and seconds of audio to look at. Every observation is metered, and the reconstruction is scored against known truth. Understanding, with a receipt.

RESEARCH PREVIEW · RECORDED DEVELOPMENT TRACES · PUBLIC LAUNCH PENDING VALIDATION

OBSERVATION / RECONSTRUCTION0:00.00

A question directs the next observation.

get_meta0 tokens
Video → selective tools → evidenceEvery observation counts.
BUILT WITH
BITTENSORQWEN2.5-VLPYTORCHTRANSFORMERSFASTER-WHISPERFASTAPIOPENCVFFMPEGBITTENSORQWEN2.5-VLPYTORCHTRANSFORMERSFASTER-WHISPERFASTAPIOPENCVFFMPEG

01 · RECORDED TRACE / ADAPTIVE POLICY

Thirty-three questions
about twenty-two seconds of video.

Every tool call, timestamp and cost below is the recorded trace of one development attempt by a 7B observer. The footage is illustrative, so you can see what each request would have captured.

RECORDED TRACE · ADAPTIVE POLICYQWEN7B · WITNESS16 ADAPTER · 33 CALLS0:00.00
get_meta
OBSERVING
VISUAL TOKENS0
AUDIO HEARD0.0 s
TRANSCRIPT0 ch
CALLS0 / 33

WHAT THE AGENT SAW

Timings and costs: recorded trace, scene 21101, benchmark 1.6 candidate. Footage: Rapid City Library, Wikimedia Commons, CC BY 3.0, used for illustration only.

02 · TWO POLICIES · ONE SCENE · SAME MODEL

Looking at everything is
the expensive way to miss it.

The same observer ran twice on the same development scene. One policy swept cheaply and zoomed where it mattered. The other requested every second at full resolution. Replayed from the recorded receipts.

ADAPTIVE POLICY160×90 SWEEP → 64×36 BURSTS → 320×180 LOOKS → 1.5 S AUDIO WINDOWS
Visual tokens
0
Tool calls
0
Audio heard
0 s
Reconstruction quality
FULL-FRAME POLICYEVERY SECOND AT 640×360 → ALL THE AUDIO
Visual tokens
0
Tool calls
0
Audio heard
0 s
Reconstruction quality

Bars scale to the larger spend. Quality is revealed when both receipts close.

03 · SEVEN EVIDENCE FAMILIES

One video.
Seven kinds of truth.

A reconstruction is not a caption. It is a timeline the scorer can check family by family, each with its own matching rule and tolerance.

  • Shot boundariesframe-accurate cuts
  • Eventswho did what, when
  • Dialoguescripted lines, timed
  • On-screen textexact strings at exact times
  • Audio eventsalarms, slams, beeps
  • Intentional errorsplanted continuity breaks
  • Question answeringgrounded in the timeline
shot boundaryevent · bee landsdialogueon-screen textaudio · wingbeatintentional errorqa · grounded
ILLUSTRATIVE OVERLAY ON CC BY 3.0 NATURE FOOTAGE (MICHAEL PISANO) · NOT MODEL OUTPUT

04 · THE METER

What does
a look cost?

Every image the validator serves is billed in 14×14-pixel patches, the units a vision transformer consumes. Resolution, frame rate and span multiply. Compose a request.

0 visual tokens

05 · EVIDENCE

Dashboard
coming soon.

Receipts, standings and the scene-level evidence behind every score will open here once independent evaluation is complete. The recorded traces on this page are a preview of that view.

DASHBOARD · COMING SOON

WHAT THE DASHBOARD WILL SHOW

  • ReceiptsEvery tool call, cost and score behind each attempt, replayable frame by frame.
  • StandingsMiners ranked by gated quality and observation efficiency, per round.
  • ScenesDevelopment scenes with the questions asked and the evidence returned.
  • ContractsThe versioned scoring rules, so a number can always be traced back to a rule.
Read the scoring contract ↗