September 03, 05:39

Google and Meta release rival AI models hours apart

Google and Meta Released Rival AI Models Hours Apart: Who Leads?

Beincrypto

Google and Meta released rival frontier AI models within hours of each other on Wednesday. Google released Gemini 3.8 Flash and a cybersecurity variant. Meta released Muse Spark 1.3. Independent testing by Artificial Analysis gave Meta an edge in agentic knowledge work and scientific reasoning. Google led in factual recall and terminal coding. Gemini 3.8 Flash is Google's third Flash release in six weeks. The model costs $0.75 per 1 million input tokens. The model costs $3.75 per 1 million output tokens. The introductory rates run through December 31, 2026. The prices then increase to $1.50 per 1 million input tokens. The output price then increases to $7.50 per 1 million tokens. Google made Gemini 3.8 Flash Cyber available to trusted defenders. The variant scored 86.2% on CyberGym. The variant scored 47.2% on CWE-Bench. Google said the model produced 2.6 times more correct patches for Chrome vulnerabilities than larger commercial models. The Fairwind Program limits access to government authorities and critical infrastructure operators. OpenAI set a similar boundary around Astra one day earlier. Meta released Muse Spark 1.3 through Muse Code and the Meta Model API. Company engineers measured roughly 20% fewer tool calls than version 1.2. Muse Spark 1.3 in max mode scored 1,754 Elo on GDPval-AA v2. Gemini 3.8 Flash in high mode scored 1,545. Muse Spark 1.3 led the Sierra Research banking agent test. Muse Spark 1.3 scored 52.4% on that test. Gemini 3.8 Flash scored 44.9%. Meta also led CritPt physics reasoning. Gemini 3.8 Flash led Terminal-Bench 2.1 with 87.6%. Gemini 3.8 Flash led AA-LCR long context with 81%. Gemini 3.8 Flash led AA-Omniscience accuracy with 55%. Gemini 3.8 Flash recorded the highest GPQA Diamond score among the tested models at 95%. The models finished within a point of each other on Humanity's Last Exam. Meta said max reasoning for Muse Spark 1.3 will arrive after further safety testing. The xhigh variant is currently available. The xhigh variant scored 61 on the Artificial Analysis Intelligence Index. That score was four points above Muse Spark 1.2. The score trailed Claude Fable 5.1 at 66. The score also trailed Claude Opus 5 at 63. Elon Musk has said Grok 4.7 will arrive shortly.

This content is an AI-generated summary/analysis for informational purposes only and does not constitute investment advice.