---
name: video-dna
description: Decompose a complete original spoken or sung video into evidence-bound scenes, shots, within-shot changes, words, sound, cast and problem progression. Produce a reusable source DNA dossier and review report. Analysis only; does not generate or adapt media.
---

# Video DNA

One original → one reusable decomposition. This is the sole video-decomposition skill
in this project, replacing `reference-narrative-dna` and `atomic-video-analysis`.
Song and dialogue are modes of the same evidence model; their timing drivers differ.
Do not turn a request to inspect a reference into production or a new product script.

## Establish what is actually in the source

Freeze original SHA-256, probe, native frame PTS, video end and audio end separately.
Inspect the complete picture **and sound**, through the commercial ending. Record the
actual model inputs/responses and sampling limits. Existing notes locate evidence;
reconcile them with source AV before accepting their claims. A frame-count check
establishes decoded coverage, not semantic coverage or creative quality.

Read [the decomposition contract](references/decomposition.md). Derive these distinct
units from observed changes: consequential scenes, editorial shots (including inserts
and returns), transition spans, events inside continuous shots, and word occurrences.
No scene, line, lyric or transport window substitutes for the shots within it.

Use [the local workflow](references/workflow.md) for the source-only tools. The native
scan measures every frame; semantic review uses actual AV plus dense PTS-labelled
sheets. Adjudicate proposed cuts on original adjacent frames, and soft transitions on
pre/inside/post context. Preserve a hold as a hold. Do not fill quiet shots with
invented gestures. Recheck brief inserts and all cuts missed by the detector.

## Join observation and causal construction

Each shot needs observed entry/exit, composition, gaze/addressee, hands/body/props,
camera, light/colour, original text, speech/delivery, sound, action/reaction events,
continuity and incoming/outgoing editorial decisions. Each event has source time,
evidence, timing precision, uncertainty and a **separate** interpretation. Missing or
unseen is valid evidence. Split concurrent channels rather than pretending every
change is a new shot.

Use raw ASR as a word index; store heard corrections separately. Never classify a
J/L cut from ASR overlap alone: check actual sound around the picture boundary.
For speech, preserve turns, breath, silence, listener changes and reaction delay.
For sung sources, add phrase/beat/arrangement/stem relationships on this same clock.
Stem leakage, tempo/key estimates and an ambiguous background remain labelled.
A quiet track is not proof that an instrument is physically absent.

Derive source passports for every recurring or consequential role, relationships,
voice, costume variants, locations, props, visual/edit grammar and problem progression.
Separate the depicted activity/result from claimed causation or medical efficacy.
Do not infer nationality from appearance; source dialogue can establish a stated
background. Do not omit a secondary person because they have few lines.

Only after the ledger, infer incoming states → witnessed event → changed states /
questions → editorial/sound choice → later dependency. Test that reading against
contrasting real source scenes and exceptions. A listener's reassurance, a buyer's
objection and a spouse's remaining doubt can require different treatments. Avoid
universal scene counts, duration, reveal ratios or mandatory gestures.

## Finish as a reusable source artifact

Freeze shot packets with exact ranges, original video, original sound, entry/middle/
exit images, dense sheets, word IDs, observations and SHA-256 manifests. Keep transition
context and source claims attached. Use half-open native-frame intervals for indexing;
soft spans remain separate and may overlap the indexed units. Do not drop speech in
a dissolve or extend the final video shot through an audio-only tail.

Show the whole original, synchronized shots/events/words/sound, scene causality,
passports, corrections and explicit limitations in a browsable report. Validate source
identity, native and semantic coverage, all evidence links, packet bytes and playback.
Report software, transport and audiovisual-reading checks separately. Match any owner
benchmark by its actual fields/evidence, not by copying a shot count or a perfect score.

Record the dossier location and scope in durable project notes. Reuse is preparation,
not permission to generate. Future production consumes this source DNA and its actual
reference packets under that project's current rules. For historical callers only,
the existing story compiler and [adaptation contract](references/call-contract.md)
remain available; they are not steps of decomposition.

Observation, interpretation, intended future output and measured output are distinct.
The owner's “10/10” is a preference, never measured retention, conversion or efficacy.
