How Soma reads a video ad.
Soma reads how an average viewer’s brain responds to a video ad, second by second, straight from the file. This page walks through how that works — the ground truth it is built on, the model that powers it, and the arc you read.
Retention graphs tell you where. Not why.
Your analytics show the drop-off. At second 14, half the viewers are gone. What they cannot show is the reason. Was it the cut, the audio, the pacing, or a promise the ad set up and never paid off?
Surveys do not recover it either. People are poor witnesses to their own attention, so asking “what did you find engaging?” returns a story assembled after the fact, not what happened while they watched. To get at the why, you have to look at the response itself, as it happens. For a long time that meant a scanner and a lab.
fMRI is the ground truth for brain response.
An fMRI scanner tracks blood oxygen across the brain. When a region works harder, it pulls in more oxygen, and the scan picks that up (the BOLD signal). Show someone a video in the scanner and you get a map of which regions responded, how strongly, and when. Not what they reported afterward. What their brain did during the ad.
Running real fMRI is slow and costly, which is why this stayed inside research labs. Scanner time runs into the hundreds of dollars an hour, and reading the data takes training. A model that predicts the response changes what is possible.
TRIBE v2 predicts that response without a scanner.
In 2025, Meta released TRIBE — the model that won the Algonauts 2025 brain-encoding challenge, 1st of 263 teams — and then TRIBE v2, the larger successor we run. Give it a video and it predicts the cortical response that video would produce, at one-second resolution across about 20,000 points on the cortical surface. It learned the mapping from real 3-tesla fMRI recordings of people watching video.
The model is public and the science is reproducible — anyone can download the same weights. Soma builds its read-out on top, turning a research-grade encoder into a product that reads ads.
The strongest signal sits below the surface.
TRIBE reads the cortical surface — the outer sheet of the brain, where attention and language load show up, and most of what a video does to you. But the strongest single neural predictor of whether an ad moves the market — the ventral striatum, the brain’s reward center — sits below that surface, where a cortical model cannot reach it.
In a head-to-head test of six methods against real ad sales, fMRI activity in the ventral striatum was the strongest predictor of market-level response — ahead of surveys, eye tracking, biometrics, and EEG. A raw cortical read is blind to it.
That does not make the cortical read empty; cortical value signals predict campaign response too. It means the surface alone is not enough. So Soma is building a layer that maps the cortical signal it can read onto the outcomes advertisers actually get — learned from a growing library of real ads paired with their real performance. It is in training, not finished, and we will say so until it is. That mapping is what a raw model — or a competitor running TRIBE straight — does not have.
Venkatraman et al. (2015), Journal of Marketing Research 52(4). Falk et al. (2015), Social Cognitive & Affective Neuroscience 11(2).
Three steps, from video to feeling.
Everything Soma shows runs through the same three steps. Here is how a video becomes an arc you can read.
Step one, video to activation, is TRIBE, benchmarked against real scans. Step two, activation to numbers, distills the activation map into a handful of signals per second. Step three, numbers to attention and feeling, is Soma’s read — the arc you see in the demo.
A few shapes come up again and again.
Once the arc is on screen, the same handful of patterns show up across ads. Here is how Soma reads them.
- see
- Attention rises early and holds through the open.
- read
- The first seconds landed and bought more time.
- see
- A spike, then a fall back to baseline before the hook resolves.
- read
- The opening got noticed but did not hold.
- see
- The arc dims where a promised payoff should land.
- read
- A slow stretch. This is the weak-spot the demo pins to the timeline.
- see
- Valence and arousal stay near neutral through a beat meant to land.
- read
- The moment is not moving anyone. The feeling reads flat.
- see
- A late lift in feeling heading into the call to action.
- read
- The ending is paying off. Feeling lifts into the close.
Where Soma goes next.
This is the order we climb, each release reading the ad more deeply than the last.
Encoder + attention arcyou are here
TRIBE runs the video; Soma reads a transparent attention arc off it, second by second.
Attention, learned
Train our own read-out head on public attention data. The first weights that are ours.
Two-dimensional feeling
The same approach reads valence and arousal — the two dimensions of feeling.
Outcomes, and the data flywheel
Retrain on partner ads paired with real audience reactions and retention. This is where Soma predicts where you lose people.
Named emotions
Amusement, tension, warmth — each named emotion read straight from the ad.