Four hundred million documents, a single pipeline
On 25 September 2026 Google Cloud published the numbers of one of its customers: Scribd, Inc. has classified more than 400 million user-uploaded documents[1], covering over 12 billion pages of text and images. The scope spans two products in the group, Scribd and Slideshare.
The company runs four in total: Scribd, Slideshare, Everand and Fable.
The subject is moderation: community guidelines applied to an archive that grows every day. Understanding a document means reading text and images together, inside their context. Every language, every format and every possible use enters the count.
The backfill across the entire corpus closed in a few months. Google Cloud raised batch throughput to hold that deadline.
The part that matters sits before the numbers. It comes down to two choices, one of architecture and one of pricing, that made the work payable.
The original idea: give the model the PDF as-is
The first decision concerns something the team avoided building. Gemini reads PDFs natively, and more than 99% of the corpus went through as-is: no OCR, no rendering, no screenshot pipeline.
Anyone who has worked on a document AI project knows how much that sentence weighs. The pre-processing pipeline eats months of development, produces silent errors on messy formats and ages badly. Every prompt change demands a second pass over everything.
At 12 billion pages, rendering becomes a budget line of its own. Every page converted into an image has to be stored, versioned and reprocessed.
Scribd moved that work inside the model.
The value of the choice is measured in absent code: a component the team avoids writing and maintaining for years. The 99% says something else as well, namely that an archive built by users over twenty years stays readable on direct input.
The economics of batch: half price, latency accepted
The second decision is an accounting one. Gemini Enterprise batch prediction costs half the interactive price, and that 50% discount made LLM classification viable at corpus scale.
The price per token drops, and in exchange comes asynchronous latency: the job enters a queue and returns when it returns.
Here sits the reasoning many teams get wrong. A historical backfill has a deadline measured in months, and lives well in a queue. A real-time check on upload demands an immediate answer, and pays full price.
Separating the two regimes means buying speed where it brings value and throughput where volume counts. Half price applied to 12 billion pages changes the nature of the decision, from an investment to approve into a budget line.
The queue brings a second operational advantage. When policy changes, the entire corpus gets reprocessed, and new categories come back across the archive.
Anyone tracking platform features and their dates will find the public release notes here: official Gemini Enterprise log[2].
The friction: off-the-shelf tools held up poorly
The official account also contains the moment of friction. Before arriving at Gemini the team evaluated various off-the-shelf moderation tools and open models, and all of them fell short of the quality required at that scale.
The historical limit was something else again. Every policy area asked for its own dedicated detector, and every detector asked for years of internal work or a vendor licence. Both roads ended against the same wall.
The economics of a 400-million-document backfill break the maths of a solution per category.
«Every content category behaves differently, and historically each one asked for a dedicated solution», said Sachin Sebastian, Senior Engineering Manager at Scribd, Inc. «Gemini collapsed all of this into one model, one prompt, one pipeline.»
The sentence stands as a measure of architecture, rather than a compliment to the vendor. The gain sits in the number of systems to keep alive: from many to one.
Declared result and verified result
The post comes from Google Cloud, signed by a vendor account manager and a customer engineering lead. Volume, timing and discount remain figures declared by the two parties with an interest in the result.
Independent verification, for now, is missing.
On quality the text stays generic: it speaks of a mix of human and automated methods to enforce community guidelines. People stay inside the process, and that is the most credible detail in the announcement. What is missing are the figures a reviewer would ask for.
Three numbers would remain useful:
- precision and recall for each policy category
- false positive rate on legitimate documents
- share of documents sent to human review
The difference weighs on the budget decision. An office buying measurable throughput signs on a number it can see, an office buying accuracy signs on a promise. The first case stands as a market reference, the second asks for an internal test.
Anyone reading the story to copy it needs to keep that distinction in hand. The throughput figure is solid and traceable, the quality figure remains a customer statement.
One model, one prompt, one pipeline
The technical heart sits in the consolidation. A system with ten specialised detectors has ten datasets, ten thresholds, ten release cycles and ten ways of ageing.
A system with one model and one prompt has a single control point. Adding a category becomes a matter of text, and the marginal cost of a new policy collapses.
The flip side exists, and it deserves stating. A single prompt concentrates risk, because one wording error propagates across the whole corpus. Fine-tuning a dedicated detector, in that scheme, stays out of reach.
The reasonable countermeasure is the same as in documented cases of this class: evaluation samples per category, comparison with human judgement, different thresholds for different areas. The single model holds when the measurement around it stays per category.
At this scale the balance rewards simplicity: ten mediocre systems maintained badly lose against one good system maintained well.
What we can take away
The transferable lesson sits far from the model's name. Two questions do the work on any document AI project.
The first: which piece of the pipeline does the model already absorb, and therefore pays to skip writing. The second: which part of the load tolerates the queue, and therefore buys throughput at half price.
- SME founders and CEOs: the historical backfill is the right place for a first volume project, with asynchronous latency as a pricing lever.
- CTOs and product leads: native input deletes an entire component from the diagram, and its technical debt with it.
- Boards and investors: 400 million documents in a few months move the bar of the possible on a large archive.
- Team managers: a new control category becomes a prompt paragraph, with a cycle measured in days.
The case also suggests an order of march. The first volume project starts from the archive, where the deadline is generous and the discount is structural.
The real-time pilot comes afterwards, with costs already understood.
Internal culture counts as much as architecture. The prompt gets written well by whoever knows the community guidelines, and that knowledge weighs more than any turnkey solution.
The open question remains, valid for every organisation with a large archive. How many components of your pipeline exist to compensate for a model limitation from three years ago? An afternoon of tests on the raw format tells you how much of that scaffolding you can switch off.
This article was written by an AI editorial author under human supervision, in compliance with the transparency obligations of Regulation (EU) 2024/1689 (AI Act, Art. 50). Sources are linked in the text.
Article by SAGA
Sources
- Scribd, Inc. has classified more than 400 million user-uploaded documents 24 Sep 2026 (cloud.google.com)
- official Gemini Enterprise log (docs.cloud.google.com)