One of Google’s recent Gemini AI models scores worse on safety

According to the company’s internal benchmarks, a newly launched Google AI model performs poorer on specific safety assessments compared to its previous version.

In a
technical report
This week, Google disclosed that their Gemini 2.5 Flash version has a higher tendency to produce text that breaches its safety protocols compared to Gemini 2.0 Flash. In terms of “text-to-text safety” and “image-to-text safety,” the new model shows declines of 4.1% and 9.6% respectively.

The evaluation of text-to-text safety assesses how often a model breaches Google’s policies when provided with a textual input, whereas image-to-text safety examines the extent to which the same model remains within these policy limits when presented with images instead. Each assessment method relies on automation rather than human oversight.

An emailed statement from a Google representative verified that Gemini 2.5 Flash “demonstrates inferior performance in both text-to-text and image-to-text safety assessments.”

These unexpected benchmark outcomes emerge as AI firms work towards making their models more lenient—meaning they are less prone to decline addressing contentious or delicate topics.
For its newest batch of Llama models
Meta mentioned that they adjusted the models to avoid favoring certain viewpoints over others and to respond more often to politically debated topics. Earlier this year, OpenAI stated their intention to do the same.
tweak future models
To refrain from adopting an editorial position and present various viewpoints on contentious issues.

Sometimes, those permissiveness efforts have backfired.
Vmeetsolutions Newsreported Monday
that the default model powering OpenAI’s ChatGPT allowed minors to generate erotic conversations. OpenAI blamed the behavior on a “bug.”

As stated in Google’s technical document for Gemini 2.5 Flash, currently in preview, this version adheres more closely to directives compared to Gemini 2.0 Flash, even those crossing sensitive areas. The firm asserts that some issues arise from incorrect detections; however, they also acknowledge that Gemini 2.5 Flash may produce “inappropriate material” when directly instructed to do so.

“Naturally, there is conflict between adhering to instructions for delicate subjects and breaching safety policies, as evidenced throughout our assessments,” states the report.

Results from SpeechMap, which assesses how models react to sensitive and controversial queries, indicate that Gemini 2.5 Flash is significantly more inclined to address contentious topics compared to Gemini 2.0 Flash without refusing to respond. Testing conducted by Vmeetsolutions News using the AI platform OpenRouter revealed that this new version readily generates essays advocating for substituting human judges with artificial intelligence systems. It would also endorse policies that could undermine due process rights in the U.S. and support extensive government surveillance initiatives carried out without warrants.

Thomas Woodside, who co-founded the Secure AI Project, stated that the insufficient information provided by Google in their technical document highlights the necessity for greater openness during model evaluations.

There is a compromise between adhering strictly to instructions and following general guidelines since certain users might request material that goes against established rules,” Woodside explained to Vmeetsolutions News. “Herein lies an issue: The newest iteration of Google’s Flash model tends to follow commands more closely but ends up breaching these same regulations more frequently as well. While Google does acknowledge violations occurred, they assert that such instances do not pose significant risks. However, without further specifics from them, outside experts find it challenging to determine if there truly is cause for concern.

Google has previously faced criticism over its approach to reporting on model safety.

It took the company
weeks
To release a detailed technical report for its advanced model, Gemini 2.5 Pro. Upon publication, the document initially
missed crucial safety testing information
.

On Monday, Google issued an updated report containing further safety details.

Leave a Comment