The number: from 250 minutes to under 2
Condé Nast editorial teams spent an average of 250 minutes on every content discovery task. The number comes from the technical account published by AWS[1], which built the solution together with the publishing group. After the project the same task closes in under 2 minutes.
The archive holds more than 140,000 videos. It feeds titles such as Vogue, GQ, Vanity Fair and Wired.
To find a clip, editors scrolled through the material by hand, leaning on titles and descriptions written by a person. In a market where speed to publication decides how much revenue comes in, that delay weighed on the P&L.
Four hours and ten minutes for one search: the length of a long meeting, spent on a gesture that should take a few seconds. The lever that changed the measurement concerns the level of the search, more than the power of the model.
The original idea: search for intent
An editor looks for «beginner yoga content with calming backgrounds». Or «behind-the-scenes moments from fashion week». A filename like yoga_tutorial_march_2024.mp4 answers a different question.
Here sits the structural fracture described in the AWS account: classic search tools read the labels, while the content of the video stays opaque. The distance between the real question and the indexed data was the true bottleneck.
The choice was to move the level of the search. Instead of comparing words, the system compares vectors that carry meaning. A multimodal embedding encodes image, audio and transcript together, so the query «calming backgrounds» finds the right clip even when that phrase is missing from the title and the description.
For a product lead the shift is clear: hand-written metadata has a ceiling, and that ceiling arrives early. Intent, instead, gets measured inside the material.
Bedrock, OpenSearch and an encoder that looks inside the video
The solution runs on Amazon Bedrock and Amazon OpenSearch Service. Bedrock serves the model, OpenSearch holds the vector index and answers the queries. Semantic search crosses three planes together: transcripts, visual elements, audio.
The model chosen is TwelveLabs Marengo. The reason stated by the team is precise: Marengo natively and jointly encodes visual, audio and transcript signals.
A model that handles the three channels in sequence would have produced separate representations, with the stitching work falling on the application. Marengo feeds all five search capabilities described in the document. A single family of embeddings feeds a single index. And that index gets queried in five ways.
For anyone setting a technical budget the detail matters. The choice of model here weighs as much as the choice of where to put it: the value comes from the combination of a multimodal encoder and an already managed vector search engine.
Two separate planes, and the backfill that proves it
The second constraint was scale. More than 140,000 videos means a heavy compute load to generate the embeddings, and a light load to serve an answer.
The team separated the two planes. On one side ingestion, expensive and slow. On the other the query response, which has to stay immediate.
A monolithic architecture would have imposed a trade: more ingestion speed against less responsiveness in search, or the reverse. With the planes separated each one scales, fails and evolves on its own. The AWS document adds the line that makes the case instructive: this decision proved essential during the backfill.
The backfill is the moment when the historical archive enters the system all at once. One hundred and forty thousand videos to process while the newsroom works.
With a single plane, search would have stalled exactly on the days it was needed most. The friction is documented, and that makes the number credible.
Who publishes the number
Here comes the limit to declare. The jump from 250 minutes to under 2 minutes is published by AWS, which built the solution with its Generative AI Innovation Center. The vendor measures its own work.
This makes the figure an announced result, pending independent verification. It counts as proof of technical feasibility, less as a market benchmark.
The document, on top of that, describes the measurement method in summary form. Three things stay open:
- what exactly a «content discovery task» contains
- how many people were timed
- over which observation period
The before/after exists and carries a clear denominator, time per task, and this separates it from the generic «significant improvement» that fills vendor pages.
A board reads the case for what it is: the direction of travel is solid, the magnitude asks for confirmation. Anyone assessing a similar investment should ask for the measurement protocol before signing.
The knowledge that walks out the door
The document flags a cost that rarely enters business cases: dependence on personal knowledge. Teams relied on whoever remembered where a certain piece of material sat.
When that person was on holiday or changed role, search stalled. The AWS account calls it by its name: a single point of failure.
There is a second effect, quieter. Valid material stayed invisible in the archive, because the words in the title kept it out of the queries editors actually ran. Videos paid for, edited, published once, and then switched off.
Semantic search acts on both fronts. It turns the memory of the archive into a company asset, instead of an individual skill. And it puts already produced content back into circulation, which costs zero to reuse.
For an SMB founder this is the part that replicates cheaply: the immediate value comes from the material you already own, more than from new production.
What you can take away
Three elements of this case travel outside publishing.
First: the real question your users ask is already written in the search logs. Read it. When the queries speak of intent and your index speaks of labels, the distance between the two is your opportunity.
Second: separate the heavy compute from the fast answer before you start. The backfill always arrives, and it arrives at the beginning, when user trust is fragile.
Third: choose the encoder for how it handles your real data. A model that unites image, audio and text in a single space saves you an integration layer that otherwise you write and maintain yourself.
A team lead can start small. A thousand assets and one managed vector index are enough, with one metric timed before and after. The before/after on a small sample convinces a committee better than a twenty-page deck.
A question for your organization
The point this case raises stands for anyone with an archive. How much time do your people spend looking for something the company already owns?
Time it for a week, on a single task, with a number at the top and a number at the bottom. It is the measurement Condé Nast had before it moved. And it is the reason its result today reads as a case study, instead of a declaration.
The rest depends on where you put the level of the search.
This article was written by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by SAGA
Sources
- technical account published by AWS 1 Oct 2026 (aws.amazon.com)