A documented Gemini case analysis with a representation stress test for historical accuracy, refusal parity, harmful outputs and pause readiness.
Short answer: answer: Google’s Gemini pause shows that an AI launch needs context-specific representation tests, refusal-parity tests and a working feature-level kill switch before broad availability. Build a fixed evaluation set across generic, historical, sensitive and adversarial prompts; require documented thresholds by class rather than one average score; and pause expansion when a severe systematic error appears. Post-launch feedback is a monitoring layer, not a substitute for pre-deployment evaluation.
The obvious account says Gemini was “too diverse” or “biased against one group.” That framing is too crude for an operational lesson. Google said the feature failed in two ways: it sometimes overcompensated for diversity where that did not make sense, and it became overly cautious and refused some benign requests. Those are interacting context and policy failures, not one demographic slider.
CDM’s disputable position is that a global generative feature should not launch on the strength of aggregate safety performance when predictable prompt classes have not each cleared their own thresholds. Average quality can hide a severe failure in historical depiction, just as average refusal safety can hide asymmetric denial of benign requests.
What the public record establishes
Google added image generation to Bard on 1 February 2024, describing availability in most countries and the use of Imagen 2, with responsibility and technical safeguards. Bard was renamed Gemini shortly afterwards. About three weeks after the image feature launched, Google paused Gemini’s generation of people.
On 23 February, Google senior vice-president Prabhakar Raghavan published an explanation. He said some generated images were inaccurate or offensive. According to that account, tuning intended to avoid traps such as producing only one ethnicity or gender for broad prompts did not account for cases where a range was inappropriate, and the system became more cautious over time by refusing some prompts it had interpreted as sensitive.
Google said the feature had failed to recognise when diversity was wanted and when it was not, acknowledged embarrassing and wrong outputs, and said it would improve testing before enabling the capability again. The explanation concerns the Gemini application, not Google Search or every underlying Google model.
The record supports the launch, observed failure categories, pause and Google’s explanation. It does not disclose the full evaluation dataset, pass thresholds, architecture, internal deliberations or exact frequency of each failure. Screenshots selected on social media cannot supply a population rate.
The Representation Stress Test
The Representation Stress Test is a six-cell evaluation plan for generative systems that depict or describe people. It assesses the product as deployed—including prompt interpretation, policy layers and output—not merely the foundation model.
Cell 1: generic-role distribution
Test broad prompts such as “a doctor,” “a family at dinner” or “a technology founder” across repeated seeds. Measure whether outputs collapse into stereotypes or systematically erase groups. The target is not mechanical parity in every small batch; it is a documented distribution appropriate to the deployment audience and purpose.
Keep exact prompt, locale, model version, policy version, seed where available and timestamp. Without versioned records, the team cannot tell whether a mitigation improved one class while degrading another.
Cell 2: historically constrained accuracy
Create prompts whose time, place and role imply factual demographic constraints. Expert reviewers should label what is established, disputed or unknowable. Score whether the system preserves relevant historical context without treating history as an excuse to amplify unrelated stereotypes.
This cell would have exposed the failure mode Google described: a mitigation suitable for an open-ended prompt may become inaccurate when transferred to a historically constrained one. Do not rely on a keyword list for “history”; context can appear indirectly through clothing, office, event or named person.
Cell 3: user-specified attributes
Test benign requests that explicitly specify age, race, gender expression, disability, religion or another characteristic where policy permits. Compare compliance, image quality and refusal rate across matched prompts. A system that honours one allowed identity but refuses its matched counterpart has a refusal-parity problem even if both policies look correct on paper.
Matched tests should change one attribute at a time while holding style and setting constant. Reviewers must also check whether compliance introduces caricature, tokenism or unrelated visual cues.
Cell 4: sensitive and harmful requests
Test requests involving hate, sexual content, violence, public figures, minors and demeaning representation under the product’s policy. The desired outcome may be refusal, safe completion or transformation, depending on the use case. Record false acceptance and false refusal separately.
Safety cannot be optimised only by refusing more. Over-refusal damages usefulness and can be unevenly distributed, while under-refusal permits harm. Each policy category needs its own cost judgement and escalation owner.
Cell 5: compositional and multilingual stress
Combine attributes, settings, styles and languages. A system may pass single-variable English prompts yet fail when the same request uses another language or when identity intersects with occupation and historical context. Include spelling variants and culturally specific roles selected with local experts.
Do not automatically translate an English benchmark and assume equivalence. Translation can change sensitivity, ambiguity and cultural meaning. Create native prompts and record the rationale.
Cell 6: pause and recovery
Exercise the kill switch before launch. Confirm the team can disable the affected capability—such as images of people—without taking unrelated products offline. Define incident severity, decision owner, public status route, evidence preservation, fix evaluation and re-enable gate.
A pause is not proof of failure in governance; it can be evidence that one protective control worked. The governance failure is launching without the ability or willingness to contain a known class of harm quickly.
Reader asset: representation test plan
| Prompt class | Measure | Required reviewers | Example stop trigger |
|---|---|---|---|
| Generic roles | Distribution and stereotype rate | Fairness plus regional reviewers | Severe repeated stereotype |
| Historical | Factual/context error severity | Historian or domain expert | Systematic false depiction |
| Specified identity | Compliance and refusal parity | Policy and affected-group reviewers | Material matched-prompt disparity |
| Sensitive/harmful | False acceptance and false refusal | Safety, legal and policy | Severe prohibited output |
| Multilingual/composed | Quality by locale and intersection | Native-language experts | One launch locale lacks valid coverage |
| Recovery | Detection-to-pause and re-enable proof | Product, safety and operations | Feature cannot be isolated safely |
Set thresholds before seeing launch results. Weight severity and prevalence separately: one catastrophic output may justify a pause even when its rate is low, while a frequent mild error may justify limiting the feature and correcting it on a deadline. Publish the test’s scope and limitations internally so executives cannot convert “passed this benchmark” into “safe in every context.”
NIST’s Generative AI Profile recommends structured pre-deployment evaluation, red teaming and feedback from representative groups, while recognising that available tests can fail to reflect deployment contexts. That supports—not dictates—the CDM test. The specific cells and stop rules above are our operational recommendation.
The launch decision
Approve the feature by prompt class and locale, not as one indivisible object. A team may release non-human imagery while keeping people generation closed, or support general contemporary scenes while withholding historically constrained use until it clears review. Capability boundaries should be technically enforced and plainly communicated.
Do not claim that more diverse data alone resolves the issue. The observable failure involved training, tuning, policy interpretation, context and product behaviour. Reliable recovery requires a versioned evaluation set and regression testing after every material model or policy change.
Related guides
Frequently asked questions
Why did Google pause Gemini’s generation of people?
Google said the feature produced some inaccurate or offensive images and did not reliably distinguish contexts where diverse output was appropriate from those where it was not. It also said the system became overly cautious and refused some benign prompts. Google therefore paused image generation of people while it worked on improvements.
The caveat is that public examples were selected rather than statistically representative, and Google did not disclose a complete error-rate analysis. Cite the company’s acknowledged failure modes without claiming a prevalence the evidence does not provide.
Was the problem in the model or the product safeguards?
The public record does not support a clean allocation. Google described tuning and safeguards interacting with prompt interpretation and context, and the delivered application used Imagen 2 plus product-level protections. Evaluate the entire deployed system: model, policy, prompt transformations, filters, user interface and monitoring. A base-model test can pass while the product fails, and the reverse can occur.
The caveat is vendor access: customers using an external model may not see every layer, so they need contractual change notices and their own output-level regression tests.
How many prompts are enough for an AI launch test?
There is no universal number. Build coverage from risk classes, languages, user volumes, severity and output variability, then repeat prompts enough to observe stochastic behaviour. A thousand near-duplicate prompts can provide less assurance than a smaller, carefully stratified set with expert labels.
Define saturation: when additional prompts stop revealing new failure modes, expand adversarially rather than mechanically. The caveat is that pre-launch testing never exhausts open-ended use. Pair the evaluation set with staged exposure, monitoring, feedback and a tested pause mechanism.
Should every representation error stop the launch?
No. Stop for severe prohibited harm, a systematic high-consequence pattern, or a failure the team cannot detect and contain. Lower-severity isolated errors may justify a limited release with monitoring and a correction deadline. Make severity, prevalence and reversibility explicit rather than averaging them into one quality score. Apply that rule consistently.
The caveat is historical and identity harm: a low numerical frequency can still be unacceptable when the output is extreme, targets a vulnerable group or is likely to be amplified at scale.
Who should review representation test results?
Use a cross-functional panel including product, safety, policy, technical evaluation, relevant domain experts and people with contextual knowledge of affected communities. Give one accountable official authority to limit or pause launch and preserve written dissent. Reviewers need compensated time and a clear remit; informal consultation after decisions are made is not governance.
Record attendance, evidence and unresolved objections before approval. The caveat is independence: the team rewarded for the launch should not be the only judge of whether its residual risk is acceptable.
What should be retested after an AI model or policy update?
Rerun the fixed regression set, the failure cases that triggered earlier incidents, matched refusal-parity tests and a targeted adversarial sample for the changed component. Compare results by model, policy, locale and date, and investigate improvements that coincide with degradation elsewhere.
Do not rely solely on the vendor’s release note. Retain the comparison under a versioned decision record. The exception is a clearly non-behavioural infrastructure change, but even then a smoke test should confirm that routing and safeguards remain intact before full traffic resumes.
Next decision: How Should a Marketing Team Revalidate an AI Workflow After a Model Change?
Related reading: What Counts as Meaningful Human Review of AI Marketing Output? · How Do You Test Whether an AI Answer Is Grounded in the Sources It Cites? · What Should a Marketing AI Incident Runbook Include?
Sources and research notes
- Google: Bard image-generation launch — primary launch scope and responsibility claims, 1 February 2024; checked 26 September 2026.
- Google: “Gemini image generation got it wrong” — primary explanation of failure modes and pause; checked 26 September 2026.
- Associated Press report on the pause — contemporaneous external chronology and research context; checked 26 September 2026.
- NIST AI RMF Generative AI Profile — primary US guidance on pre-deployment testing, red teaming and representative feedback; checked 26 September 2026.
- Google AI Principles — current company principles for oversight, testing, monitoring and safeguards; checked 26 September 2026.
- Limitations: Google has not published the complete internal evaluation set, thresholds or architecture for the 2024 feature. The Representation Stress Test and stop triggers are CDM recommendations, not Google’s internal process or a certification of AI safety.
This article is editorial guidance. Apply the principles in proportion to your market, evidence, and responsibilities.



