🤖 China Just Dropped the Largest Open AI Model Ever — and That's Only the Start

Welcome to Daily Inference, your daily dose of the most important developments in artificial intelligence. I'm your host, and today is Friday, July 18th, 2026. We've got a packed show covering frontier model releases, enterprise AI growing pains, a fascinating fashion trend born from surveillance anxiety, and OpenAI playing offense against its own systems. Let's get into it.

But first, a quick word from today's sponsor, 60sec.site. Need a website fast? 60sec.site uses AI to build you a stunning, professional website in under sixty seconds. No coding, no design skills needed. Check them out at 60sec.site.

Alright, let's start with a massive model drop out of China. Moonshot AI released Kimi K3 on July 16th, and this thing is enormous. We're talking 2.8 trillion parameters — making it the largest open AI model to come out of China. Now, that number sounds astronomical, but here's the clever part: the model uses a Mixture-of-Experts architecture, meaning it only activates 16 out of 896 experts at any given time. So despite its staggering size, it's actually efficient in practice. It also supports a one-million token context window, which means it can process and reason over incredibly long documents in a single pass. The Rundown AI is calling it a frontier gap closer, and based on early reports suggesting it could rival Anthropic's Opus 4.8, that framing seems warranted. This is a clear signal that the open-source model race is no longer a Western-dominated affair — China is competing at the very top of the capability stack.

Speaking of model releases, NVIDIA dropped something important on the retrieval side of AI. The company released Nemotron 3 Embed on July 15th and 16th, a collection of open embedding models. The flagship 8-billion parameter version has already claimed the number one spot on the RTEB benchmark — that's the Retrieval Text Embedding Benchmark — with a score of 78.46. What's interesting here is the engineering efficiency story. NVIDIA also released a smaller one-billion parameter version that was created through a process called pruning and distillation from the larger model, retaining nearly all the retrieval quality at a fraction of the compute. And the NVFP4 variant runs at up to twice the throughput on NVIDIA's Blackwell hardware. These aren't just benchmark wins — they're building blocks for the kind of enterprise retrieval infrastructure that's desperately needed right now, which brings us perfectly to our next story.

VentureBeat published a sweeping series of enterprise AI research this week, and the picture it paints is one of remarkable ambition running well ahead of operational reality. Across multiple surveys covering hundreds of enterprises, a consistent theme emerges: companies are deploying AI faster than they can measure, secure, or trust it.

Start with infrastructure. Eighty-three percent of enterprises report their GPUs are running at fifty percent utilization or less, yet they're planning to spend heavily on specialized AI cloud infrastructure they barely use today. Fewer than half can even rigorously track what their AI compute actually costs. They're buying the next layer of hardware before they've figured out the economics of what they already own.

Then there's the agent security picture. More than half of enterprises surveyed — fifty-four percent — have already experienced either a confirmed AI agent security incident or a near-miss. And yet only a third give each AI agent its own dedicated credentials. Most agents are sharing login information, which means a single compromised agent can potentially act with far more reach than intended. Companies are largely relying on security guardrails bundled in from their AI providers like OpenAI and Google, rather than purpose-built agent security tools.

And the evaluation problem might be the most unsettling finding of all. Half of organizations have shipped an AI agent that passed their internal testing, then failed a customer in production. Only five percent say they fully trust automated evaluations. Yet two-thirds are already allowing, or actively building toward, fully automated deployments with no human in the loop. The autonomy is racing ahead of the assurance. It's a pattern we'll be watching closely.

Now let's talk about something that sits at the intersection of AI safety and something OpenAI is doing that's genuinely fascinating. The company has built an internal AI model called GPT-Red — essentially an AI hacker that it uses to attack its own systems. Trained using self-play reinforcement learning, GPT-Red beat human red-teamers eighty-four percent to thirteen percent at identifying prompt injection vulnerabilities. It also discovered a novel attack class called Fake Chain-of-Thought, where it manipulates an AI's reasoning process to produce unsafe outputs. And when used to harden GPT-5.6, it reduced failures on OpenAI's hardest direct injection benchmark by a factor of six. This is a significant milestone in AI safety methodology — using AI offensively to systematically close security gaps before bad actors find them. Though OpenAI acknowledges it still struggles with multi-turn and image-based attacks, the approach represents a meaningful evolution beyond human-only red teaming.

Now let's pivot to something you might wear to your next protest. Fashion designers in the UK are developing what they're calling adversarial clothing — garments embedded with carefully designed patterns of shapes, colors, and repeating motifs that are engineered to confuse facial recognition systems. As facial recognition gets rolled out across more public spaces in Britain, a new generation of designers is turning privacy into a fashion statement. Think of it as cybersecurity for your wardrobe. The patterns exploit known weaknesses in computer vision models — the same underlying vulnerabilities that researchers have been documenting for years in academic settings, now translated into streetwear. Whether this becomes a mainstream movement or stays niche is an open question, but the fact that it's a conversation happening at all speaks volumes about where public anxiety around AI surveillance is headed.

Finally, let's touch on one of the more chilling stories in the AI space this week. A humanoid robotics company called Foundation Future Industries, which counts Eric Trump as its chief strategy adviser, has told WIRED it's actively exploring military applications for its robots — what the CEO diplomatically described as exploring some, quote, kinetic things. This comes as the Dutch navy is simultaneously testing autonomous uncrewed vessels for naval defense, and the broader defense robotics sector continues to accelerate. We're watching a world where AI-powered physical systems are moving from warehouse logistics to potential combat roles, and the ethical and regulatory frameworks simply haven't kept pace.

That's your Daily Inference for today. The through-line across all of these stories is a world moving faster than it can govern itself — whether that's enterprises deploying agents without proper security, regulators scrambling to keep up with model capabilities, or society beginning to wear its anxieties about surveillance literally on its sleeve.

For more analysis and the stories we couldn't fit into today's show, head to dailyinference.com and subscribe to our daily AI newsletter. We break down the most important AI news every single day so you never miss a beat. Thanks for listening, and we'll see you tomorrow.

🤖 China Just Dropped the Largest Open AI Model Ever — and That's Only the Start
Broadcast by