🤖 Grok 4.5 Is Here & It's Cheaper Than You Think — Plus OpenAI's Voice Mode Just Changed Everything
Welcome to Daily Inference, your go-to source for the most important developments in artificial intelligence. I'm glad you're here, because today's news cycle is absolutely packed. Let's get into it.
But first, a quick word from our sponsor. If you've ever wanted to build a website but didn't want to spend hours wrestling with code or design tools, check out 60sec.site. It's an AI-powered platform that lets you create a beautiful, functional website in under sixty seconds. Seriously. Head over to 60sec.site and see for yourself.
Alright, let's talk AI.
We're kicking things off with a major model drop from SpaceXAI. Grok 4.5 is now live, and Elon Musk is calling it an Opus-class model — which is a direct comparison to Anthropic's most powerful Claude tier. What makes this release particularly interesting is that it was trained in collaboration with Cursor, the popular AI coding assistant. The result is a model purpose-built for coding, agentic workflows, and knowledge-intensive tasks. It runs at 80 tokens per second, and here's where it gets really competitive — pricing comes in at just two dollars per million input tokens. For context, that's significantly cheaper than many frontier alternatives. And it's already topped Harvey's Legal Agent Benchmark, which evaluates how well AI handles complex legal reasoning. So this isn't just a coding tool — it's positioning itself as a serious contender across professional domains. The Cursor partnership is a smart move too. Rather than building a general-purpose model and hoping it sticks, SpaceXAI went straight to where developers actually work and baked Grok into that workflow from the ground up.
Speaking of voice and interaction, OpenAI dropped something quietly significant this week. They released GPT-Live and a companion model called GPT-Live-1 mini — a completely overhauled voice mode for ChatGPT. What's different here is the architecture. These are full-duplex models, meaning they can listen and speak simultaneously, just like a real phone call rather than a walkie-talkie. The model is also designed to interrupt you less, and if you pause mid-thought, it actually waits instead of jumping in. When it needs to do heavy lifting — like searching the web or working through complex reasoning — it hands off to GPT-5.5 in the background and then comes back to continue the conversation. OpenAI's research team is calling GPT-Live-1 their smartest voice model yet. This has real implications for live translation, accessibility tools, and ambient AI assistants. The era of awkward, stilted AI voice conversations may genuinely be behind us.
Now let's talk about efficiency, because NVIDIA just made a quietly impressive engineering move. They released a new model called Nemotron-Labs-3-Puzzle-75B-A9B — yes, that's a mouthful. But here's what it actually means in plain terms. NVIDIA took their larger Nemotron-3-Super model, which had over 120 billion total parameters, and used a technique called Iterative Puzzle to compress it down significantly — alternating between hardware-aware compression and short recovery training sessions so the model doesn't lose too much capability. The result is a model with about 75 billion total parameters that delivers more than twice the server throughput of the original on a single eight-GPU node. Even more striking — on a single H100 chip, it can handle eight simultaneous million-token context requests where before it could only handle one. This is the kind of infrastructure-level work that quietly determines which AI products are actually economically viable at scale. Doubling throughput without doubling hardware costs is a big deal for anyone deploying AI in production.
On the robotics front, Ant Group's Robbyant lab just open-sourced LingBot-VLA 2.0, and this one is worth paying attention to. It's a six-billion-parameter vision-language-action model — meaning it sees the world, understands language, and then takes physical actions. What's remarkable is the training data behind it: roughly 60,000 hours total, including 50,000 hours of robot trajectory data across 20 different robot configurations, plus 10,000 hours of video from humans doing everyday tasks in first-person view. The model maps all of these different robot body types into a single unified action space covering arms, hands, waists, heads, and mobile bases. That's the cross-embodiment breakthrough — one model that can control dramatically different physical robots. It outperforms competing models including pi-zero-point-five on benchmark testing, and it's released under Apache 2.0, meaning anyone can use it commercially. This is exactly the kind of open-source push that could accelerate the robotics industry the same way open LLMs accelerated the software AI space.
And that brings us to a broader trend worth connecting here. A Bezos-backed startup called General Intuition is making a bold bet that video game data — millions of hours of it — could be the missing ingredient for achieving artificial general intelligence. The argument is compelling: current language models are exceptional at text but genuinely struggle with understanding how objects move through physical space and time. Video games, on the other hand, are rich simulations of physics, causality, and spatial reasoning. If you can train on that data at scale, you might get AI that understands the world the way robots need to understand it to function. Pair that insight with what Robbyant is doing with real robot trajectories, and you start to see a picture of how physical AI might actually mature — not through one magic breakthrough, but through layering diverse, high-quality training signals from multiple domains.
Before we wrap up, let's touch on something that cuts across all of these stories. The AI buildout is accelerating — more models, more products, more funding. SambaNova just raised a billion dollars at an eleven-billion-dollar valuation. Lovable is reportedly in talks to double its valuation to over thirteen billion. Prime Intellect pulled in 130 million dollars to help enterprises train their own agents. The capital keeps flowing. But there's a shadow side to this. Datacenter expansion is straining energy grids and water supplies — Meta's Wyoming datacenter had a contractor flush bacteria-contaminated water into public sewers during construction. Google and Amazon are watching their net-zero climate commitments slip further out of reach. Communities from London's Brick Lane to towns across Australia are pushing back. The AI industry is going to have to reckon with this tension between scale and sustainability, and probably sooner than most leaders want to admit.
That's your Daily Inference for today. Lots happening, and we're tracking all of it. For deeper dives and daily updates delivered straight to your inbox, head over to dailyinference.com and subscribe to our newsletter. We break down the most important AI stories every single day so you don't have to scroll through the noise. Thanks for listening — we'll see you tomorrow.