The science

What Soma measures, and what it doesn't.

Soma predicts how an average viewer's brain responds to a video ad, second by second, straight from the file. This page walks through how that works, which part is proven, and which part we are still testing. We label both, because the line between them is the point.

01 · the gap

Retention graphs tell you where. Not why.

Your analytics show the drop-off. At second 14, half the viewers are gone. What they cannot show is the reason. Was it the cut, the audio, the pacing, or a promise the ad set up and never paid off?

Surveys do not recover it either. People are poor witnesses to their own attention, so asking "what did you find engaging?" returns a story assembled after the fact, not what happened while they watched. To get at the why, you have to look at the response itself, as it happens. For a long time that meant a scanner and a lab.

02 · ground truth

fMRI is the ground truth for brain response.

An fMRI scanner tracks blood oxygen across the brain. When a region works harder, it pulls in more oxygen, and the scan picks that up (the BOLD signal). Show someone a video in the scanner and you get a map of which regions responded, how strongly, and when. Not what they reported afterward. What their brain did during the ad.

time → BOLD stimulus peak response, a few seconds later
The blood-oxygen response lags the stimulus by a few seconds and then settles. TRIBE learns this signal from real scans.

Running real fMRI is slow and costly, which is why this stayed inside research labs. Scanner time runs into the hundreds of dollars an hour, and reading the data takes training. A model that predicts the response changes what is possible.

03 · the model

TRIBE v2 predicts that response without a scanner.

In 2025, Meta released TRIBE — the model that won the Algonauts 2025 brain-encoding challenge, 1st of 263 teams — and then TRIBE v2, the larger successor we run. Give it a video and it predicts the cortical response that video would produce, at one-second resolution across about 20,000 points on the cortical surface. It learned the mapping from real 3-tesla fMRI recordings of people watching video.

The model is public. Anyone can download the same weights. That cuts two ways for us: the science is reproducible, and the model by itself is not a moat for anyone. We build the read-out on top of it, and we say plainly that the encoder is Meta's, not ours.

About that “92%”

video → activation: validated

You'll see a “92% correlation with fMRI” figure quoted by other brain-AI tools and attributed to Meta. We can't source it to any Meta publication, so we don't use it. Meta's TRIBE actually reports a mean correlation of about 0.21 across ~1,000 cortical regions on held-out data — roughly half of the measurable ceiling — with TRIBE v2 several-fold better again. Either way, that number is the accuracy of video → brain activation: it says nothing about whether activation predicts whether someone keeps watching. That next step is a separate question, and it is the one we test in the open.

04 · the chain

Three steps. Only the first is proven.

Everything Soma shows runs through the same three steps. Pulling them apart is the honest way to look at the product, because the three do not carry the same weight of evidence.

Video ad the file you upload
validated TRIBE v2,
vs real fMRI
Brain activation ~20k points · 1 Hz
descriptive no claim
attached
Summary numbers a few signals / sec
validating our read-out,
tested vs humans
Attention & feeling the arc you read

Step one, video to activation, is TRIBE. Benchmarked against real scans. Not ours. Step two, activation to numbers, is arithmetic over the activation map, descriptive with no claim attached. Step three, numbers to attention and feeling, is our read. It is a hypothesis, and we are validating it now against real human data. When the demo labels a line "attention," it is this third step, and it wears an amber badge for a reason.

05 · the trap

A spike is not a feeling.

The most common mistake in this field is reading a burst of activation as one specific emotion. The same peak could be interest, confusion, mild alarm, or noise. Activation on its own cannot tell you which. Only behavior settles it: did they keep watching? We will not claim a feeling from activation alone, and you should be wary of anyone who does.

06 · how we test it

We test the third step in the open.

The question is whether the predicted arc tracks a real human attention curve. We test it inside a single video, second against second, on public data where the human answer already exists: TVSum, where 20 people rated how interesting each shot was. We line our predicted arc up against theirs.

clip time → range expected by chance (permutation null) Soma predicted human
Illustration of the test, not a result — the measured result from our first run (n = 15) is stated just below.

A few guardrails keep the test honest. We compare the shape of the change second to second, not the slow drift, so a lucky trend cannot pass as signal. We shuffle the timing thousands of times to see what a random arc would score, and we only count what beats that. We check the predicted arc against a plain baseline of loudness, cuts, brightness, and motion, so we can tell whether the brain read adds anything over a dumb feature detector. And we write the plan down before we look, so a null result gets reported the same as a positive one.

The uncomfortable part, from us

disclosed prior

A published result found that whole-brain activation does not predict which parts of a YouTube video get replayed. So the whole-cortex version of our signal is a likely dead end, and we treat it as the baseline to beat. The narrower, region-specific test is the one still open. We would rather you hear that from us than find it in diligence.

What the first run actually showed

validating · early · n=15

We ran it on 15 TVSum clips. The raw arc came back null, exactly as pre-registered — and it does not beat the loudness/cuts/brightness/motion baseline (0 of 15 clips). The raw arithmetic arc, on its own, is not the product.

But a small trained read-out over the region features does track the human interest curve, held out video by video: median rank correlation r ≈ 0.20 — about 87% of the agreement humans reach with each other — combined p = 0.0003. And that signal survives the loudness/cuts/brightness/motion control (partial r ≈ 0.18, p = 0.0005), so it is not just re-deriving the edit.

Honest limits, stated plainly: 15 videos, so any single clip is underpowered; TVSum measures interest — a public proxy, not ad retention; and the read-out is a learned hypothesis, not a validated engagement model. That is exactly how we report it. (Frozen snapshot of the first n=15 run.)

The next test: does it rank real ads?

cross-sectional · not yet run

TVSum measures interest inside one clip. The question a marketer actually asks is different: given a set of real ads that really ran, does our score pick the winner? So we run the same honest yardstick across ads — take ~15–30 ads with a known real outcome (a partner’s own CPA / ThruPlay, or a public proxy like TikTok Top-Ads rank), score each one, and check whether our score ranks them by performance after removing loudness, cuts, and length — so we can’t win just by re-detecting “short and loud.” Pre-registered primary, permutation null, effect floor, and the null gets reported like any other.

This is the test that turns “we can predict your winning ad” from a hope into a number — and until that number clears the bar, we don’t make the claim. (Result: pending the first real-ad run.)

07 · how to read the arc

A few shapes come up again and again.

Once the arc is on screen, the same handful of patterns show up across ads. Here is how we read them today. Each is a working interpretation we are still testing, not a settled rule, so each carries its evidence tier and a note on what would prove it wrong.

1 2 3 4 5
  1. 1Visual — where the ad enters the brain
  2. 2Motion — cuts and camera movement
  3. 3Faces & voices — the biology wedge
  4. 4Attention — does the ad hold
  5. 5Feeling — valence and arousal
Illustrative regional view, not a per-vertex render. The read-out on the demo is bound to model output; this drawing is for orientation.
Strong hookvalidating
see
Attention rises early and holds through the open.
read
The first seconds landed and bought more time.
Wrong if: ads with this shape drop off as fast as ads without it.
Weak hookvalidating
see
A spike, then a fall back to baseline before the hook resolves.
read
The opening got noticed but did not hold.
Wrong if: these ads retain as well as ones that sustain the rise.
Mid-video leakvalidating
see
The arc dims where a promised payoff should land.
read
A slow stretch. This is the weak-spot the demo pins to the timeline.
Wrong if: the flagged second is not where viewers actually leave.
Flat feelinghypothesis
see
Valence and arousal stay near neutral through a beat meant to land.
read
The moment may not be moving anyone. The affect read is an unproven proxy.
Wrong if: the proxy fails to track LIRIS human affect ratings.
Strong closehypothesis
see
A late lift in feeling heading into the call to action.
read
The ending may be paying off. Same caveat: affect is a proxy, not a decoder.
Wrong if: the lift does not line up with real end-of-video response.
08 · the roadmap

We add a claim only when a test reproduces it.

Each rung is earned by a held-out test, not asserted. This is the order we climb, and we do not skip ahead.

R0

Frozen encoder + honest attention arc you are here

TRIBE runs the video; we read a transparent attention arc off it and badge it as a hypothesis under test.

R1

Attention, learned and validated

Train our own read-out head on public attention data, checked leave-one-video-out. The first weights that are honestly ours.

R2

Two-dimensional feeling

The same approach reads valence and arousal against public human-labeled data, every row badged as a proxy.

R3

Outcomes, and the data flywheel

Retrain on partner ads paired with real audience reactions and retention. This is where "predicts where you lose people" earns its claim.

R4

Named emotions

Amusement, tension, warmth, each one earned by a held-out test, never before.

See it on a real ad How Soma compares Read the FAQ