Welcome to The Median, DataCamp’s newsletter for July 10, 2026.
In this edition: OpenAI releases GPT-5.6 and GPT-Live while preparing to sunset the Atlas browser, Meta introduces Muse Spark 1.1 alongside new media generation models, SpaceXAI launches Grok 4.5, China considers restricting foreign access to its top AI models, and the Future of Life Institute publishes its Summer 2026 AI Safety Index.
This Week in 60 Seconds
OpenAI Releases GPT-5.6 and Introduces GPT-Live
OpenAI has released the GPT-5.6 model family and GPT-Live, a new framework for voice interaction. The GPT-5.6 release introduces three distinct capability tiers: Sol, Terra, and Luna. Sol is engineered for complex tasks requiring high-level coding, cybersecurity, and scientific reasoning. Terra is balanced for everyday professional workflows, while Luna is optimized for fast, cost-efficient workflows. Performance metrics indicate that Sol is at least on par with Anthropic’s Fable 5, with OpenAI’s self-reports positioning it above the competitor on several coding and agentic benchmarks while operating at a significantly lower cost. OpenAI also released GPT-Live, an update to the ChatGPT Voice experience that allows the system to speak and listen concurrently to mimic natural conversational rhythms.
Meta Releases Muse Spark 1.1 and New Media Models
Meta Superintelligence Labs has expanded its product ecosystem this week with the rollout of the Muse Spark 1.1 reasoning model and its new media generation platform. Muse Spark 1.1 is built for personal agentic tasks, enterprise-grade software engineering, and multi-app planning across a context window of 1 million tokens. The company also launched Muse Image and introduced a preview of Muse Video to handle text-to-video capabilities with native audio support. Muse Image functions agentically by executing web searches and code routines to self-refine its visual output, while also featuring an invisible watermarking layer called Content Seal to aid in provenance verification. We’ve covered Muse Spark 1.1 and the new media models on our blog.
SpaceXAI Launches Grok 4.5
SpaceXAI has released its new flagship model, Grok 4.5. The model supports advanced software development, autonomous AI agents, and corporate knowledge workflows. Grok 4.5 falls behind Fable 5 and GPT-5.5 on most self-reported benchmarks. To optimize codebase navigation, the model was co-trained alongside the AI code editor Cursor, allowing it to interpret developer interactions and navigate multi-file projects. Beyond engineering, the system functions as an office assistant capable of conducting web research to create financial models, native presentation layouts, and legal documents. The model is currently available to U.S. developers and users via Cursor, Grok Build, and the SpaceXAI API Console, with availability expanding to the European Union later this month.
OpenAI Sunsets Atlas Browser to Focus on ChatGPT Work
OpenAI has announced the discontinuation of its standalone Atlas browser, framing the project as a learning experiment in agentic web browsing. To prioritize the ChatGPT Work suite launched this week, the company is unbundling the browser’s core technology and integrating it directly into its existing software ecosystem. Moving forward, users will access these capabilities through a built-in browser within an upgraded ChatGPT desktop application, alongside a new Google Chrome side-chat extension. This move follows the closure of Sora, an AI video-generation social media app, earlier this year. As we noted in a previous newsletter, the decision to close Sora was part of a broader pre-IPO consolidation strategy, and the sunsetting of Atlas likely falls under the same initiative.
China Considers Restricting Foreign Access to Its Top AI Models
According to an exclusive report by Reuters, Chinese authorities have held meetings with top technology firms, including Alibaba, ByteDance, and Z.ai, to discuss potentially restricting overseas access to China’s most advanced AI models. Officials also raised the possibility of implementing new regulatory measures to restrict who can fund domestic AI startups. The proposed limits under discussion could apply to both closed-source and open-weight architectures, mirroring recent restrictions implemented by the US over its own frontier systems due to national security concerns. Given the growing global adoption of efficient Chinese models by international businesses, market analysts suggest that any eventual export curbs could shift operational costs across the global AI market.
Future of Life Institute Releases Summer 2026 AI Safety Index
The Future of Life Institute has published its Summer 2026 AI Safety Index, an independent scorecard ranking nine major AI developers across six core safety and governance domains. Evaluated by a panel of technical experts using the US GPA grading system, the report shows that Anthropic, OpenAI, and Google DeepMind lead the industry, though no model developer scored higher than a C+. Conversely, firms such as xAI, DeepSeek, and Mistral received failing marks. The index highlights that amid intense market competition, progress in commercial safety has largely plateaued or regressed, with multiple labs scaling back prior safety commitments to pursue national security and defense contracts. We will explore this topic in more depth in the Deeper Look section below.
Agentic Data Science & Engineering (July 20-24)
Learn how Microsoft, Stanford, Accenture, and top practitioners are using agents to reshape data analysis, science, and engineering workflows.
A Deeper Look at this Week’s News: The AI Safety Index (Summer 2026 Edition)
Ever since the introduction of ChatGPT, AI regulation has remained a constant social and political debate globally. We have seen the introduction of complex regulatory frameworks like the EU AI Act and, most recently, the US government blocking foreign access to a model deemed capable of exploiting software vulnerabilities.
The current landscape suggests we’ll see more and more AI regulation in the coming months and years. AI companies can and need to improve their safety practices, particularly as model capabilities accelerate. In this landscape, the Future of Life Institute published its biannual AI Safety Index, ranking nine major AI companies across key safety and security domains.
What is the AI Safety Index?
The AI Safety Index, published twice a year by the Future of Life Institute, offers an independent look at how the AI industry approaches safety. By evaluating the top AI companies against a consistent set of metrics, the report shows how safety and security practices evolve over time.
An independent panel of seven technical and governance experts grades the evidence, which is gathered from public records, model documents, and industry surveys.
The grading uses the US GPA system (ranging from A+ to F) and evaluates model developers across six areas: risk assessment, current harms, safety frameworks, existential safety, governance, and information sharing.
The AI safety leaderboard
The summer 2026 scorecard shows a clear separation at the top, with Anthropic, OpenAI, and Google DeepMind leading by a significant margin.
However, the baseline for success remains remarkably low. No model developer scored above a C+. This suggests that poor safety practices are a global issue rather than a regional one. The bottom tier also reflects the global aspect, with three failing grades spread across three different continents: xAI in North America, DeepSeek in Asia, and Mistral in Europe.
Mistral’s performance is especially interesting because it illustrates what researchers call “European dissonance.” While the European Union is a global leader in AI regulation, its most prominent homegrown AI startup ranked last with a score of just 0.33.
Beneath the rankings, the expert panel identified a disconnect between public messaging and commercial conduct, noting that corporate safety frameworks often lack quantitative risk thresholds and independent audits.
Under competitive pressure, major developers have weakened their prior commitments to unilaterally pause development, shifting instead to conditional postures contingent on competitors’ behavior.
The weakening of these commitments coincides with an industry shift toward defense contracts, as leading labs have reversed prior bans on military applications to pursue national security agreements.
How AI safety metrics evolved over time
A look at the scorecard over the past three evaluation cycles indicates an industry struggling to sustain long-term safety progress. The capability of AI models continues to accelerate, but safety metrics have largely plateaued or slipped backward.
Anthropic has remained static, moving from 2.64 last summer to 2.66 today. Both OpenAI and Google DeepMind followed a familiar arc, peaking in late 2025 before sliding back in the current report. Translating written safety frameworks into concrete operational practice is proving difficult for the market leaders.
Meta represents the sole example of steady improvement, climbing incrementally from 1.06 in Summer 2025 to 1.32 today. Conversely, xAI has experienced a severe collapse, dropping from a stable D grade throughout 2025 to a failing 0.65 this summer.
This volatility seems common among the trailing labs. DeepSeek temporarily rose to 1.02 in late 2025 before crashing to 0.47. Coupled with Z.ai’s slide to 0.88 and Mistral’s low debut at 0.33, the numbers indicate that the safety and security gap between the market leaders and the rest of the industry is widening.
To explore the AI Safety Index in greater detail, check the following reports:
Industry Use Cases
Stanford Scientists Build an AI Lab Partner
A team of scholars at Stanford University has developed Biomni, an open-source, general-purpose biomedical AI agent designed to execute complex research tasks alongside human scientists. The platform integrates large language models with over 150 bioinformatics tools, 59 databases, and 100 software packages to automate workflows like data analysis, molecular modeling, and experimental design. During its initial nine months of operation, more than 15,000 researchers used the tool to execute over 100,000 distinct scientific workflows, including extracting patterns from raw genomic sequences and continuous glucose monitor datasets. Read more in this article from Stanford HAI.
Mistral AI Introduces Robostral Navigate
Mistral AI has launched Robostral Navigate, an 8B-parameter model built for autonomous robotic navigation in complex indoor and outdoor environments using only a single RGB camera. Unlike traditional multi-sensor setups that depend on LiDAR or depth sensors, the model processes standard visual images alongside plain-language instructions to guide wheeled, legged, or flying mechanisms. In benchmark testing, the system achieved a 76.6% success rate on unseen Room-to-Room in Continuous Environments (R2R-CE) validation datasets, outperforming existing single-camera and multi-camera alternatives. Read more in this article from Mistral AI.
Government of Alberta Uses Claude to Address Cybersecurity Vulnerabilities
The Government of Alberta, through its Ministry of Technology and Innovation, has integrated Anthropic’s Claude Code to identify and remediate cybersecurity vulnerabilities across its provincial digital infrastructure. Using Claude Opus and Sonnet, an internal team successfully scanned 466 million lines of code across 3,400 software repositories in approximately 20 hours. Beyond identifying hidden security flaws, the autonomous agents generated corresponding code patches, developed automated testing suites, and successfully modernized outdated applications, including a 25-year-old public subsidy portal. Read more in this article from Anthropic.
Tokens of Wisdom
The worst thing about AI is the lies that have no tells — when a system tells you something that’s false, and it looks just like an answer that’s true.
—Dan Klein, CTO at Scaled Cognition
This week on the DataFramed podcast, we met Dan Klein, CTO at Scaled Cognition. We discussed why AI reliability has lagged behind capability, how hallucinations hide in plain sight, the limits of humans-in-the-loop, and more.





