OpenAI's GPT-6 Sol and Luna: cheaper on paper, but nerds aren't so sure
OpenAI says its new models are a bargain - but early testers running real workloads are finding the maths doesn't always add up.
Artificial intelligence, models, tools and their consequences.
OpenAI says its new models are a bargain - but early testers running real workloads are finding the maths doesn't always add up.
A new open-source web agent promises to click through websites in seconds flat, and developers are piling in — though the eye-catching benchmark comes entirely from its own demo.
Hacker News is buzzing about Xiaomi's openness in training a new AI model — but 'transparent' isn't the same as 'independently verified', and the benchmarks are still Xiaomi's own.
A new open-source agent orchestrator has gone viral on Hacker News under a Google banner, though the project's own site never quite says so.
A new open-source browser agent claims a flight search in 7.1 seconds, but the maths behind that number, and the reason for its meteoric GitHub popularity, deserve a second look.
A new preprint says robot-controlling language models ignore safety instructions almost every time — and offers a fix that still fails one job in three.
A browser-automation project from browser-use raced up GitHub's charts in under two weeks, though the numbers behind the hype deserve a second look.
A new Commission wants Britain to be the best place to build, use and trust medical AI, but for now this is advice to ministers, not a change to what happens in your GP surgery.
A one-page announcement and a sprawling Hacker News argument about ChatGPT's name is, for now, most of what the internet actually knows about Meta's new assistant.
Ministers say firms will be freed from red tape and 'pen-pushing', with AI even name-checked as a future fix - but the actual rules, dates and detail are conspicuously absent.
OpenAI has shipped GPT-6 Astra with headline-grabbing benchmark scores, but a look at the small print - and a sceptical Hacker News crowd - suggests the comparisons doing the AGI talk aren't quite apples-to-apples.
An 'autonomous research system' has gone from nothing to 6,000-plus GitHub stars in under a fortnight, yet the documentation trails off mid-sentence and nobody's checked whether it actually works.
Claude Fable 5.1 is cheaper and generally available; its sibling Mythos 5.1, with fewer safety brakes, is locked behind vetting that most people will never see.
A new statistical tool from the Allen Institute for AI says popular LLM benchmarks quietly mix up different skills into one score, but the checking is being done by the people who built the method.
A research team has released the full design and simulator for a cable-driven robotic hand, claiming policies trained purely in software work on the real thing without adjustment - a bold claim that, for now, rests entirely on a preprint.
Headlines say Nvidia has snapped up the open-source AI hub for $13 billion, but the actual reporting says the deal isn't done yet.
A TypeScript project from the DeepSeek stable has rocketed up GitHub's charts in under a fortnight, though the details of what it does, and whether the numbers mean what they seem to, remain thin on the ground.
A TypeScript project from DeepSeek's GitHub org has rocketed to six-figure star counts in under a fortnight, yet the actual pitch amounts to a tagline and a folder full of AI agent scaffolding.
A TypeScript project from DeepSeek's GitHub org has rocketed up the charts, but the numbers tell you more about hype than about whether the thing actually works.
The company that lets developers shop between hundreds of AI models says nothing will change for users, but the reported price tag and the motive behind the deal are worth a second look.
Alibaba's latest small local model produces some of the best output yet seen from a laptop-sized LLM, but only after it's finished agonising over the question.
Google shipped a new "workhorse" model three weeks after the last one. The coding numbers are all first-party, the launch price is introductory, and UK users cannot reach it in the Gemini app at all yet.
Sponsored results reached British users on 11 August. They appear on the Free and Go tiers, they are chosen from what you type, and paying more is the only way to switch them off.
Kassper puts a person on every plan. Two years ago the AI SDR category sold itself on removing exactly that, and the retreat says more about the market than the launch does.
Inevitable AI Group has raised a $6m pre-seed to mass-produce software companies. The plan is striking, the budget per venture is thin, and one figure in the announcement does not match the record.