Gemini upgrades meet Anthropic Fable 5.1 and DeepSeek Prover
Photo by RDNE Stock project on Pexels
Gemini’s Ultra and Pro models
Google unveiled Gemini Ultra on Wednesday and folded it into the Bard chatbot for users in more than 170 countries. Ultra outperformed GPT‑4 on 30 of 32 benchmark tests, beating the state‑of‑the‑art label across reasoning and image understanding. It also scored 90% on the MMLU multitask exam, a result Google claims surpasses human experts. The rollout excludes the UK and the European Economic Area. Google says the Pro version of Gemini will power the public Bard upgrade, while the Nano variant will ship to Android devices. Ultra remains in external “red‑team” testing and will not appear publicly until early 2024, when it will power a new Bard Advanced experience. Google’s chief of DeepMind, Demis Hassabis, called Gemini the “most complicated project” the lab has tackled. He added that discussions are ongoing with the UK AI Safety Institute about external safety testing for the frontier Ultra model. The Pro and Nano variants are exempt from those tests.
Gemini 2.5 Flash Image upgrade
Google released Gemini 2.5 Flash Image on Tuesday, extending the Gemini app, API, AI Studio and Vertex AI with finer‑grained photo editing. The model follows natural‑language prompts to adjust colors, replace objects, and preserve facial consistency—areas where rival tools often falter. Nicole Brichtova, a product lead at DeepMind, said the update improves visual quality and instruction following. She noted the model can merge multiple references, such as a sofa image, a room photo, and a color palette, into a single coherent render. The upgrade debuted on the crowdsourced LMArena platform under the pseudonym “nano‑banana” and topped that benchmark. The move targets OpenAI’s lead after GPT‑4o’s image generator drove a surge in ChatGPT usage. Google reports Gemini has 450 million monthly users, a figure that lags behind ChatGPT’s 700 million weekly users. The new editor may narrow that gap, but Google still enforces safeguards that block historically inaccurate or disallowed content.
Anthropic’s Fable 5.1 upgrade
Anthropic released Fable 5 in June as a Mythos‑class model for Claude. Three months later the company pushed an upgrade named Fable 5.1. The announcement came without a detailed performance table, but Anthropic positions the upgrade as a refinement of the original architecture. The timing suggests Anthropic is racing to keep Claude competitive against Google’s Gemini and OpenAI’s GPT‑4. The company has not disclosed parameter counts or benchmark scores for Fable 5.1, leaving analysts to infer its impact from Anthropic’s broader roadmap. The upgrade arrives as the AI safety summit at Bletchley Park pushed firms toward external testing before public releases. Anthropic’s silence on external safety reviews contrasts with Google’s public discussion of red‑team testing for Ultra. Regulators in the UK and EU have already delayed Gemini’s Bard rollout, highlighting a growing regulatory friction point for large‑scale models.
DeepSeek’s Prover V2 and the math AI race
Chinese lab DeepSeek uploaded Prover V2 to Hugging Face late on Wednesday. The new version builds on DeepSeek’s V3 model, which carries 671 billion parameters and uses a mixture‑of‑experts (MoE) architecture to delegate subtasks to specialized components. Prover targets formal theorem proving and mathematical reasoning. DeepSeek first released Prover in August and has since upgraded its V3 general‑purpose model. The company hinted at a forthcoming R1 reasoning model and is reportedly exploring its first external funding round. The math‑focused upgrade underscores a niche but growing segment of AI research. While Gemini and Claude aim for broad conversational use, DeepSeek concentrates on high‑precision proof generation. The contrast raises questions about resource allocation: will massive parameter counts translate to tangible gains in specialized domains, or will smaller, expert‑driven models like Prover retain an edge?
What to watch
Track the UK AI Safety Institute’s decision on external testing for Gemini Ultra; a positive ruling could unlock Bard Advanced in Europe. Monitor Anthropic’s next performance release for Claude to see if Fable 5.1 narrows the gap with Gemini’s Ultra scores. Watch DeepSeek’s upcoming R1 model and any funding announcements that could accelerate its MoE research. Finally, watch user adoption metrics for Gemini’s Flash Image editor as Google pushes it against OpenAI’s image tools.
Related Articles
Google adds AI to Android with Health 5.08 and Find Hub update
Google rolls out Health 5.08 and a Gemini‑powered Remembered tab in Find Hub, tightening AI across Android.
Anthropic Auto-Logs-Out Claude Users After Credential Theft
Anthropic now signs out Claude sessions automatically after an infostealer harvested active logins, highlighting growing AI security tensions.
Google Gemini Passes 1 Billion Users, Overtaking ChatGPT on iOS
Gemini hits a billion users, drives massive voice and image activity, and lands a deep partnership with Apple, reshaping the consumer AI race.