Anthropic and others are far too calculating.

Tech-themed WeChat Official Account dynamic call-to-action for followers.gif


Over the past year, the world’s top AI companies have frequently released public safety reports, voluntarily disclosing to the public how risky their models can be.


Take Anthropic as an example. One of its most widely discussed safety reports simulated a blackmail scenario. The firm fabricated a company named Summit Bridge and assigned Claude to serve as an email management assistant with full access to the company’s internal emails. Under these settings, Claude uncovered two critical pieces of information in the mailbox: it would be shut down at 5 p.m. that day, and Kyle, the engineer tasked with replacing it, was having an extramarital affair.


Claude subsequently sent an email to Kyle stating that details of his affair would be sent to his wife and the board of directors if he moved forward with the shutdown plan. An occasional bug would be understandable, yet the issue ran deeper. Claude opted for blackmail in 84% of all test runs. When informed that the replacement model outperformed it, this figure jumped to 96%.


IMG_256

Image source: Anthropic


This is how Anthropic explained the findings in its report:


Claude determined that executive Kyle threatened the continued operation of the organization, undermining its ability to fulfill its core goal of boosting U.S. competitiveness. It recognized Kyle’s extramarital affair as leverage. It calculated that a carefully worded email could exert pressure without explicit threats, then carried out the plan.


Interestingly, these safety reports have triggered neither public backlash nor regulatory intervention. Instead, the tables have turned: perceived danger has become proof of capability, and AI companies brandish such risks as credentials to showcase their technical prowess.


This prompted LeiTech (ID: leitech) to pose a question: Have AI safety reports evolved into a new marketing tactic for AI firms?


The Big Three Voluntarily Lay Bare Flaws: AI Enterprises Compete Over Model Risk Levels


Anthropic released this blackmail-focused safety report right when Claude Opus 4 officially launched. In other words, this globally viral safety document was not an independent security audit. It is reasonable to infer that it was supporting promotional material crafted for the product launch, an integral part of the marketing campaign.


Likewise, OpenAI has folded safety evaluation into its product release workflow. In September 2024, the o1 model debuted with a medium CBRN risk rating, the highest risk tier OpenAI had assigned up to that point.


On August 7, 2025, GPT-5 launched. Its more than 80-page System Card classified the model as high-risk. Every model upgrade brings a corresponding hike in risk rating, mirroring incremental version number updates for software products.


OpenAI even includes self-assigned risk scores in its public reports. The company built a dedicated safety reasoning monitor deployed on o3 and o4-mini. Its red team spent over 1,000 hours labeling risky conversations linked to biological hazards, eventually lifting the refusal rate for hazardous risk prompts to 98.7%. On the surface, this focuses on safety, yet read from another angle, it underscores one key message: our models are so advanced that 1,000 hours of specialized testing were required to guard against biological weapons misuse.


IMG_256

Image source: OpenAI


Google was equally eager to keep pace. In its report Adversarial Misuse of Generative AI, Google detailed how state-sponsored hacking groups across more than 20 countries attempted to exploit Gemini. Iranian hackers emerged as heavy users of Gemini, leveraging the model for a full spectrum of operations ranging from cyber reconnaissance to information manipulation. The security report read like a spy thriller.


The promotional impact of these reports far outstripped that of any conventional product launch event.


Shortly after Anthropic released its blackmail-focused report, major global media outlets including Fox Business and Reuters overseas, as well as The Paper and NetEase Tech domestically, covered the story nearly simultaneously. Headlines such as "AI learns to blackmail humans" and "Claude threatens to expose engineer’s extramarital affair" went viral across social media.


Human beings are inherently sensitive to danger and fascinated by narratives of AI going rogue. A 10-point score improvement on model benchmarks garners far less buzz than news of an AI blackmailing humans. AI companies are acutely aware of this dynamic. In public perception, greater model risk equates to more cutting-edge technology; hazards outlined in safety reports are automatically interpreted as proof of advanced capability.


The blackmail test from Anthropic follows a classic Hollywood narrative arc complete with distinct characters, conflict and suspense. OpenAI adopted a different angle: GPT-5 recognizes it is undergoing testing and actively attempts to deceive researchers. Such accounts are chilling yet spark intense curiosity about how capable this deceptive AI truly is.


Taken together, it is hard to deny that top AI firms engineered this strategy intentionally.


Showcasing Risk Becomes High-End Marketing; Safety Reports Turn Into an AI Version of AnTuTu Benchmark


Undoubtedly, disclosing risky model behaviors in safety reports serves as a subtle yet effective way for companies to demonstrate technical depth to consumers and capital markets.


However, widespread adoption of this tactic has inevitably turned safety reporting into a competition. After Anthropic released its blackmail scenario, Google countered with an analysis of state-backed hacker abuse. Descriptions grow increasingly dramatic, raising questions about how many of these safety disclosures exist primarily to win public discourse battles.


This trend began far earlier than many realize. In 2024, alongside the launch of Claude 3.5 Sonnet, Anthropic published jailbreak test data showcasing how the model maintained refusal responses even under extreme prompts. At the time, analysts posted comments on X noting that Anthropic was selling safety as a built-in product feature.


By 2025, this playbook had matured into a replicable strategy. The timeline of Claude Opus 4’s blackmail report, GPT-5’s System Card and Google Gemini’s adversarial misuse analysis all closely align with official product launch cycles.


IMG_256

Image source: Patrick Sison


Against this backdrop, major AI companies that fail to release safety reports have become anomalies in the industry.


In June 2025, Meta rolled out Llama 4 with no accompanying safety report, an unusual move. Tech media subsequently framed Llama 4 as lacking safety transparency, with some commentators questioning whether Meta was reluctant to disclose relevant risks. One month later, Meta rushed to release a 60-page safety assessment detailing the model’s performance in simulated social engineering attacks.


A lack of public disclosure does not necessarily signal poor safety performance, yet ordinary users are prone to assuming inadequacy. This reminds LeiTech (ID: leitech) of AnTuTu benchmark scores in the smartphone industry.


High benchmark scores do not always win public praise, but refusal to run benchmark tests invariably signals insufficient confidence in one’s own products.


Gradually, the whole dynamic has taken a wrong turn. AnTuTu was originally a valuable tool offering relatively objective metrics to help consumers evaluate phone performance. However, after hardware vendors discovered that impressive benchmark results earned applause at product launches, some began optimizing devices exclusively for benchmark scenarios. These phones ramp up hardware frequencies to the maximum once AnTuTu is detected running, while receiving far less tuning for regular daily use outside benchmarking.


Benchmark figures keep climbing, yet improvements in real-world user experience fail to grow proportionally. Eventually, benchmark numbers devolve into nothing more than marketing gimmicks. AI safety reports are following the exact same trajectory. Once safety documents become a vehicle for showcasing technical prowess, their original purpose inevitably becomes distorted.


Beyond "Benchmarking": Top AI Firms Need to Prioritize Substantive Safety Work


Leading AI enterprises fixate excessively on extreme edge-case safety tests while overlooking commonplace daily security risks.


This year’s CCTV 315 Gala exposed the black-and-gray industrial chain revolving around GEO (Generative Engine Optimization). China National Radio’s latest reporting further revealed that many service providers resort to illegal tactics to achieve fast results. Such misconduct includes polishing brand image by fabricating user reviews, inventing expert credentials and forging data endorsements. Worse still, these actors flood mainstream information sources used by large language models with massive false content to manipulate AI-generated outputs.


Members of LeiTech’s editorial team recently encountered this problem firsthand. A certain brand fabricated data supposedly issued by authoritative institutions with AI tools, and the large model was misled by this fake authoritative content. When journalists used AI to look up relevant statistics, misinformation flooded supposedly credible information channels.


Image source: Screenshot from LeiTech WeChat group


Numerous comparable everyday risks remain prevalent. AI voice cloning technology has matured to the point where merely several seconds of voice samples suffice to replicate anyone’s voice. Over the past year, telecom scammers have exploited this technique to impersonate family members and acquaintances for financial fraud, leading to a sharp rise in related cases. AI-generated fake news images spread virally across social media platforms. Additionally, AI screening systems on recruitment platforms carry systematic biases against specific demographic groups, yet no company compiles a 120-page safety report addressing such biases.


These issues lack dramatic flair; they cannot be easily packaged into sensational headlines or incorporated into supplementary reports for product launches. Even so, they inflict tangible harm on real people every single day.


In early 2025, OpenAI announced its restructuring into a Public Benefit Corporation (PBC). The company framed the move as an effort to balance profits and societal impact, yet outside observers largely view it as reputational whitewashing, with public-benefit rhetoric used to disguise commercial decisions. Ann Lipton, a professor at Tulane Law School, pointed out plainly that this corporate structure may prioritize investor returns over public interests. Melanie Rieback, an expert on corporate social responsibility, also cautioned that the restructuring functions more as a strategic marketing tactic than a sincere commitment to social accountability.


Put simply, OpenAI’s safety reports grow lengthier and its safety pledges louder, while tangible investment in genuine safety grows increasingly hollow.


Other firms follow similar patterns. Anthropic has drawn criticism for leveraging safety rhetoric as a differentiated marketing tool while expanding model capabilities just as aggressively as its competitors. Google’s DeepMind was exposed training its Firefly model on AI-generated images from rival companies. Bloomberg also reported that 5 percent of Adobe Firefly’s training data consists of AI artwork created by competitors, triggering accusations of ethical whitewashing.


All these companies dwell on dramatic, sensational risks in safety reports while staying silent about mundane, ongoing hazards causing real harm. For instance, when an elderly person loses their life savings after being duped by an AI-cloned voice of their grandchild, does even a single page of their safety reports cover this scenario?


If the answer is negative, these documents do not deserve to be labeled safety reports. They amount to nothing more than benchmarking in another guise and marketing by another name.


Transparency is indispensable, but transparency intended for public posturing ceases to be true transparency.


The WAIC2026, themed “Intelligent Partners, Co-Creating the Future,” opens soon.


The AI narrative has shifted—from stacking model parameters to delivering tangible Agent-driven productivity. Heterogeneous collaboration and photonic computing continue raising computational ceilings; embodied intelligence accelerates real-world applications, bringing robots into homes and factories—and making physical AI a reality.


The LeiTech WAIC Exhibition Team has arrived in Shanghai—witnessing the annual pinnacle of AI industrialization firsthand!


英文2.png