The thesis . Chapter five . 6 min read

The data layer.
Nobody buys an asset without due diligence.

Who, what, when, and how, for every minute. The unglamorous layer that decides the quality of everything built on top of it.

Danielle Dafni
Danielle Dafni
August 4, 2026
An inspection bench under a lamp: tweezers lifting a single frame from an open film reel, cut frames and a wax seal on the desk, colored frames pinned along the front, and a tall stack of wax-sealed cans standing unopened beside it

In the first chapter of this thesis I presented the claim that video contains five layers: data, insights, assets, knowledge, and above them, the action layer. I presented them then as a list, one paragraph each, and moved on. But as the series progressed, the stock, the balance sheet, the portfolio manager, I realized I had done the layers an injustice. Each one is a world, and each deserves a chapter of its own.

So over the coming chapters we will do exactly that: break video down layer by layer, from the bottom up. And we start with the foundation. The layer everyone skips, and without which everything above it collapses.

You cannot price what you have not examined

In the world of investing there is a rule nobody argues with: before you buy an asset, you do due diligence. Nobody acquires a company based on the cover. You open the books, examine the contracts, talk to the customers, verify what is actually inside. The worst deals in history are almost always the deals where someone skipped this step.

Now look at the enterprise video library. Hundreds of hours of assets that nobody has ever examined.

Not metaphorically. Literally: nobody in the organization can say who appears in hour three of the 2024 seminar, what was claimed there, which customers were mentioned, and what was promised from the stage. The asset exists, but its contents, in any deep sense, are unknown to everyone.

This is the data layer: the due diligence of video. It answers questions that are almost embarrassingly simple, and that almost no organization can answer today:

  • Who?Who are the speakers, what are their roles, which organization are they from, who spoke with whom.
  • What?What was said, word for word, and who said it. Not "roughly." Exactly.
  • When?At which minute each topic came up, when the speaker changed, where every moment lives.
  • How?Facial expressions, body language, tone, context. The moment the audience laughed. The moment the speaker hesitated. The slide on screen behind them.

It sounds basic. It is basic. And that is exactly the point: without this foundation, nothing in the floors above it is possible.

Why a transcript alone is worth almost nothing

Here I want to take apart a common misunderstanding. When organizations hear "data layer," they think: ah, transcription. We have that. We ran an automated tool, we got a text file, box checked.

But a raw transcript is not a data layer, for the same reason a pile of numbers is not a financial statement. A financial statement is not the numbers. It is the structure: which number belongs to which line, what is revenue and what is expense, what changed from quarter to quarter. Take the same numbers, remove the line labels, and you have noise.

On the left a tangle of identical colorless film loops. On the right the same film sorted into separate colored strips, each clipped by a peg of its own color and running along its own rail to its own spool

A transcript without structure is exactly that: noise. Three hours of a panel collapsed into one block of text, where you cannot tell the moderator from the guest, a claim by the CEO from a quote he was citing, an official statement from a joke. Anyone who tries to build on it, a clip, a quote for the press, an answer to a business question, will fail. And after the first failure, the organization stops trusting the library, and the whole effort dies.

A real data layer is a transcript that is attributed and anchored: every sentence tied to a speaker, every speaker to an identity and role, every moment to a timestamp, and all of it to the visual context, what was happening on screen and in the room at that second. The difference between the two is the difference between "we have text" and "we have a documented record of what happened."

The boring layer is the layer that decides everything

Let me say something unmarketable: the data layer is the least glamorous layer in the pyramid. No CEO gets excited about speaker identification. The excitement is reserved for the upper floors, the insights, the assets, the ability to talk to your library.

But here is the truth I have learned in seven years in this field: the quality of the entire pyramid is set at the bottom layer. An insight resting on a wrong attribution is a wrong insight. A clip cut at the wrong moment is a defective asset. An answer from a knowledge system that confuses two speakers is worse than no answer, because it sounds confident.

A massive block base taking up most of the frame with a small lit stage, podium and cameras on top of it, while three figures work at the base itself with spirit levels and a plumb line

In investing, this is called garbage in, garbage out. Brilliant analysis of falsified statements is worth nothing. In content, the same rule holds exactly:

No layer can be more accurate than the layer beneath it.

This is also why the figure from the first chapter, only about 10% of unstructured information is even stored, and even less is analyzed, is not just a storage problem. Even the little that is stored is mostly stored without a data layer. Which means even the organizations that "keep everything" are holding an archive of assets that never went through due diligence. They do not know what they have. They only know that they have.

What this enables as early as tomorrow morning

And despite calling it boring, the data layer on its own, before any of the floors above it, changes things on the ground:

  • Search that works"The moment the customer from X talked about their results" goes from two hours of scrubbing to five seconds of query. This is the liquidity we discussed in the stock chapter, and it starts here.
  • Zero dependence on human memoryToday, the knowledge of "what is in that conference" lives in the head of the marketing manager who was there. When she leaves, it leaves with her. A data layer turns people's memory into an organizational asset.
  • A foundation for every future requestEvery clip, quote, speaker kit, or answer that anyone asks for a year from now is already waiting, anchored and attributed. The investment is made once. The dividends are collected forever.
A table set with a grid of round sockets, every film reel in a socket of its own color, three hands lifting three different reels at the same moment, and a tangle of unsorted film waiting at the edge

And in the language of this series: this is the IPO. The moment a private, opaque, untradable asset becomes a listed one.

The bottom line

Due diligence is not the reason you buy an asset. It is the reason you can trust everything you do with it afterward. So it is with the data layer: it is not the exciting story. It is the reason the exciting story will be credible.

An open presentation case with a wax seal in its lid and an examined film reel inside, a blank tag tied to it, the inspection tools laid down beside it, and an open hand reaching in to receive it

An organization asking "what is our video library worth" first needs to answer a more modest question: do we even know what is in it? Who, what, when, and how, for every minute?

If the answer is no, the good news is that this is the most solvable problem in the pyramid. The technology for this breakdown exists, works, and is fast. At Speechbox, it is the floor everything else is built on.

Next chapter: the insights layer. The difference between market data and an analyst's call, and why this is the layer executives actually consume.

Danielle Dafni
Danielle Dafni

Founder and CEO of Speechbox, a platform that turns enterprise video into an active knowledge asset.

Sources

  1. Research World, Possibilities and limitations of unstructured data. Only about 10% of unstructured information is stored, and less is analyzed.
  2. Box and IDC, Untapped Value white paper. The gap between what organizations hold and what they ever examine.

Questions this raises

Is a transcript a data layer?

No, for the same reason a pile of numbers is not a financial statement. A financial statement is the structure, not the figures. A real data layer is a transcript that is attributed and anchored: every sentence tied to a speaker, every speaker to an identity and role, every moment to a timestamp, and all of it to what was on screen at that second.

What does the data layer actually answer?

Who spoke, in what role and from which organization. What was said, word for word, and by whom. When each topic came up and where every moment lives. And how it was said, including expressions, tone and the slide on screen behind the speaker.

Why does the accuracy of the bottom layer matter so much?

No layer can be more accurate than the layer beneath it. An insight resting on a wrong attribution is a wrong insight, a clip cut at the wrong moment is a defective asset, and an answer that confuses two speakers is worse than no answer, because it sounds confident.

What does a data layer change immediately, before anything is built on top of it?

Search starts working, so finding the moment a customer described their results goes from two hours of scrubbing to a query. Institutional knowledge stops depending on the person who happened to be in the room. And every clip, quote or kit anyone asks for a year from now is already anchored and waiting.

How I am building this

Want to see what your own archive is worth?

At Speechbox we turn raw video into clean, scored, sellable assets. The appraisal and appreciation layer this whole idea needs.

See Speechbox

The signal, not the noise.

Every time a new deal or study proves the thesis, I send it with one line on what it means. That is the whole newsletter.