Skip to content

How Garfunkel analyzes a video

Paste a link and Garfunkel turns the video into a structured report: what’s said, what’s shown, the hook, the same set of tags as every other post and, if you ask, how people reacted. Here is each step, in order.

Copies this page as Markdown, ready to paste into an AI assistant.

  1. Step 01 of 08

    Getting the post

    We open the link and fetch the video together with what the platform shows about it: the caption, the creator, and the views, likes, comments and shares it reports. Those numbers are kept exactly as reported, so you can sort and compare posts by them later. The video file is only kept while we work on it and is deleted once the analysis is finished.

    Getting the postA video link becomes a post card with its thumbnail, creator, caption and view, like and comment counts.Link…/video/1842@creator1.2M84K1.1K
  2. Step 02 of 08

    Splitting the video

    We check whether the video has a soundtrack, then separate the sound from the picture so that each can be examined on its own. The sound goes on to be transcribed and the picture goes on to frame sampling, and the two come back together when the post is summarized.

    Splitting the videoA video splits into two: a filmstrip of the picture and a waveform of the sound.VideoPictureSound
  3. Step 03 of 08

    Transcribing the audio

    A speech model writes down everything said in the video, with a timestamp on every line, so you can see what was said and when. A video with speech is never analyzed without its transcript: if this step fails, it is retried rather than skipped, and a post that fails is never charged.

    Model: Qwen3-ASR-0.6B · an open-source speech recognition model

    See it in the Transcript tab of a result

    Transcribing the audioA waveform becomes three transcript lines, each with a timestamp.0:000:030:07Transcript0:00
    So here’s the trick…
    0:03
    you heat the pan first,
    0:07
    then add the oil.
  4. Step 04 of 08

    Choosing the frames

    We use a frame sampling algorithm to reduce the number of frames while keeping the overall meaning of the video intact. The frames that remain are what the AI model looks at, together with the transcript, in the steps that follow.

    Choosing the framesA long strip of near-identical frames is reduced to a few frames that each show something different.All framesSampled frames
  5. Step 05 of 08

    Summarizing

    An AI model reads the transcript and the caption while looking at the frames, so it understands the post as a whole rather than one piece at a time. It writes a short summary and the story in order, beat by beat, with each beat tied to a moment in the video.

    Model: gpt-6.1-sol · reads the transcript and the sampled frames together

    SummarizingTranscript lines and frames flow into a summary card that lists the story beats in order, with timestamps.Transcript0:000:030:07FramesSummaryStory in order0:00
    Sets up the trick
    0:04
    Heats the pan first
    0:09
    Shows the result
  6. Step 06 of 08

    Finding the hook

    The model looks closely at the first few seconds to name the hook: the words or image that grab attention, what kind of hook it is and the moment it lands. It also notes the promise the hook makes to the viewer, how the video pays it off and why it works.

    Model: gpt-6.1-sol · the same model and the same pass as the summary

    See the Hook card on a result

    Finding the hookA timeline of the opening seconds with the hook window highlighted, its type and quote, and rows for the promise, the payoff and why it works.Opening seconds0:000:010:020:030:04Demonstration
    “So here’s the trick…”
    PromiseFrying with nothing stuck to the pan.
    PayoffThe food slides off cleanly at the end.
    Why it worksShows the method instead of telling it.
  7. Step 07 of 08

    Tagging and analysis

    With the summary, transcript and frames, the model fills in the same set of tags for every post: format, theme, emotion, tone, people and brands, places, objects, editing, and cultural or meme references. The key labels are double-checked for consistency. Because every post gets the same set, any two posts can be compared side by side.

    Model: gpt-6-luna · fills in the tags and checks for cultural and meme references

    See the tags on a result

    Tagging and analysisA summary card feeds a grid of nine tags: format, theme, emotion, tone, people and brands, places, objects, editing and culture.Summary
    FormatTutorial
    ThemeFood & drink
    EmotionJoy
    TonePlayful
    People & brandsOne creator
    PlacesHome kitchen
    ObjectsPan, oil
    EditingJump cuts
    CultureTrending sound
  8. Step 08 of 08

    Reading the comments

    When you ask for comments, we collect them and the model reads each one in the context of the post: its caption, summary and transcript. That way a reply is judged against what the video actually said and showed, and the report shows who agrees, who pushes back, what people ask and what they liked.

    Model: gpt-6-luna · reads each comment against the post

    Add comments to a result

    Reading the commentsThree comments, each labeled agrees, pushes back or asks, above a bar showing where people stand.
    Finally, someone explained it properly.Agrees
    That’s not how it works, though.Pushes back
    Where did you get that pan?Asks
    Where people stand

How the context stacks

Each step builds on the one before it, so the later steps never work from the raw video alone. This is what each one reads.

  1. 01TranscriptReadsThe soundtrack
  2. 02Summary and hookReadsTranscript, caption and sampled frames
  3. 03TagsReadsSummary, transcript, caption and sampled frames
  4. 04Comment reportReadsComments, caption, summary and transcript

Every post, the same shape

Every post is analyzed into structured output with a fixed schema: the same fields, in the same format, for every post. That keeps results consistent and comparable, easy to sort and filter, and clean to export to a spreadsheet, the API or MCP.

{
  "format": "Tutorial",
  "theme": "Food & drink",
  "emotion": ["Joy"],
  "tone": ["Playful"],
  "hook": { "type": "Demonstration", "start": 0.0, "end": 2.0, "text": "So here’s the trick…" },
  "tags": ["home kitchen", "pan", "jump cuts"],
  "transcript": [{ "start": 0.0, "end": 2.4, "text": "So here’s the trick…" }],
  "comments": { "agree": 52, "push_back": 14, "questions": 18 }
}

The field names here are illustrative; the docs list the real ones.

What you get

  • A summary and the story in order
  • A timestamped transcript
  • The hook: what it is, when it lands and why it works
  • The same set of tags as every other post
  • A comment report, when you ask for one
  • Search by meaning: find posts by what they’re about, not only by the words in them
  • Exports to spreadsheets, the API and MCP

The video file is deleted once the analysis is finished.