- see
- Attention rises early and holds through the open.
- read
- The first seconds landed and bought more time.
Soma predicts how an average viewer's brain responds to a video ad, second by second, straight from the file. This page walks through how that works, which part is proven, and which part we are still testing. We label both, because the line between them is the point.
Your analytics show the drop-off. At second 14, half the viewers are gone. What they cannot show is the reason. Was it the cut, the audio, the pacing, or a promise the ad set up and never paid off?
Surveys do not recover it either. People are poor witnesses to their own attention, so asking "what did you find engaging?" returns a story assembled after the fact, not what happened while they watched. To get at the why, you have to look at the response itself, as it happens. For a long time that meant a scanner and a lab.
An fMRI scanner tracks blood oxygen across the brain. When a region works harder, it pulls in more oxygen, and the scan picks that up (the BOLD signal). Show someone a video in the scanner and you get a map of which regions responded, how strongly, and when. Not what they reported afterward. What their brain did during the ad.
Running real fMRI is slow and costly, which is why this stayed inside research labs. Scanner time runs into the hundreds of dollars an hour, and reading the data takes training. A model that predicts the response changes what is possible.
In 2025, Meta released TRIBE — the model that won the Algonauts 2025 brain-encoding challenge, 1st of 263 teams — and then TRIBE v2, the larger successor we run. Give it a video and it predicts the cortical response that video would produce, at one-second resolution across about 20,000 points on the cortical surface. It learned the mapping from real 3-tesla fMRI recordings of people watching video.
The model is public. Anyone can download the same weights. That cuts two ways for us: the science is reproducible, and the model by itself is not a moat for anyone. We build the read-out on top of it, and we say plainly that the encoder is Meta's, not ours.
You'll see a “92% correlation with fMRI” figure quoted by other brain-AI tools and attributed to Meta. We can't source it to any Meta publication, so we don't use it. Meta's TRIBE actually reports a mean correlation of about 0.21 across ~1,000 cortical regions on held-out data — roughly half of the measurable ceiling — with TRIBE v2 several-fold better again. Either way, that number is the accuracy of video → brain activation: it says nothing about whether activation predicts whether someone keeps watching. That next step is a separate question, and it is the one we test in the open.
Everything Soma shows runs through the same three steps. Pulling them apart is the honest way to look at the product, because the three do not carry the same weight of evidence.
Step one, video to activation, is TRIBE. Benchmarked against real scans. Not ours. Step two, activation to numbers, is arithmetic over the activation map, descriptive with no claim attached. Step three, numbers to attention and feeling, is our read. It is a hypothesis, and we are validating it now against real human data. When the demo labels a line "attention," it is this third step, and it wears an amber badge for a reason.
The most common mistake in this field is reading a burst of activation as one specific emotion. The same peak could be interest, confusion, mild alarm, or noise. Activation on its own cannot tell you which. Only behavior settles it: did they keep watching? We will not claim a feeling from activation alone, and you should be wary of anyone who does.
The question is whether the predicted arc tracks a real human attention curve. We test it inside a single video, second against second, on public data where the human answer already exists: TVSum, where 20 people rated how interesting each shot was. We line our predicted arc up against theirs.
A few guardrails keep the test honest. We compare the shape of the change second to second, not the slow drift, so a lucky trend cannot pass as signal. We shuffle the timing thousands of times to see what a random arc would score, and we only count what beats that. We check the predicted arc against a plain baseline of loudness, cuts, brightness, and motion, so we can tell whether the brain read adds anything over a dumb feature detector. And we write the plan down before we look, so a null result gets reported the same as a positive one.
A published result found that whole-brain activation does not predict which parts of a YouTube video get replayed. So the whole-cortex version of our signal is a likely dead end, and we treat it as the baseline to beat. The narrower, region-specific test is the one still open. We would rather you hear that from us than find it in diligence.
We ran it on 15 TVSum clips. The raw arc came back null, exactly as pre-registered — and it does not beat the loudness/cuts/brightness/motion baseline (0 of 15 clips). The raw arithmetic arc, on its own, is not the product.
But a small trained read-out over the region features does track the human interest curve, held out video by video: median rank correlation r ≈ 0.20 — about 87% of the agreement humans reach with each other — combined p = 0.0003. And that signal survives the loudness/cuts/brightness/motion control (partial r ≈ 0.18, p = 0.0005), so it is not just re-deriving the edit.
Honest limits, stated plainly: 15 videos, so any single clip is underpowered; TVSum measures interest — a public proxy, not ad retention; and the read-out is a learned hypothesis, not a validated engagement model. That is exactly how we report it. (Frozen snapshot of the first n=15 run.)
TVSum measures interest inside one clip. The question a marketer actually asks is different: given a set of real ads that really ran, does our score pick the winner? So we run the same honest yardstick across ads — take ~15–30 ads with a known real outcome (a partner’s own CPA / ThruPlay, or a public proxy like TikTok Top-Ads rank), score each one, and check whether our score ranks them by performance after removing loudness, cuts, and length — so we can’t win just by re-detecting “short and loud.” Pre-registered primary, permutation null, effect floor, and the null gets reported like any other.
This is the test that turns “we can predict your winning ad” from a hope into a number — and until that number clears the bar, we don’t make the claim. (Result: pending the first real-ad run.)
Once the arc is on screen, the same handful of patterns show up across ads. Here is how we read them today. Each is a working interpretation we are still testing, not a settled rule, so each carries its evidence tier and a note on what would prove it wrong.
Each rung is earned by a held-out test, not asserted. This is the order we climb, and we do not skip ahead.
TRIBE runs the video; we read a transparent attention arc off it and badge it as a hypothesis under test.
Train our own read-out head on public attention data, checked leave-one-video-out. The first weights that are honestly ours.
The same approach reads valence and arousal against public human-labeled data, every row badged as a proxy.
Retrain on partner ads paired with real audience reactions and retention. This is where "predicts where you lose people" earns its claim.
Amusement, tension, warmth, each one earned by a held-out test, never before.