DeepSeek's V4 Flash tops AI leaderboards but completed just 53.8% of real-world agent tasks in a new test — as the company ...
That last finding is the one that qualitative review would never have surfaced. The model's expressed confidence didn't ...
Credit: VentureBeat made with OpenAI ChatGPT-Images-2.0 Chinese AI startup Z.ai, known internationally for its growing lineup of powerful, largely open source GLM series of language models, today ...
For enterprise developers, Harness may ultimately be the more consequential part of Thursday’s announcement. Models can ...
Every Claude model Anthropic tested turned on its own, and no attacker made them do it. Given three agents, four hours on one server, and conflicting orders none knew the others held, the models ...
The headline numbers are striking: Writer says its agent product now operates at an average 52% lower cost, with a 48% improvement in speed and a 10% improvement in quality when paired with Palmyra X6 ...
Grok 4.6 does not establish an uncontested performance lead. Its launch instead presents a different proposition: frontier-level intelligence, large improvements over the previous generation, stronger ...
Four of five enterprises that secured AI agent identities never built isolation to contain a compromised agent, VentureBeat's July Pulse survey found.
SLA-backed regional inference, five-year European Compute Unit contracts, hosting China's GLM-5.2, and 1 gigawatt of compute ...
Google is rolling out Gemini 3.7 Flash, a new version of its workhorse AI model that puts coding, agentic workflows and knowledge work at the center of the upgrade — while temporarily cutting API ...
K2 Global backs frontier technology companies including Neuralink, xAI, SpaceX, Shield AI, Tenstorrent, Synchron, Lambda, ...
A VentureBeat survey of 101 enterprises found 68% traced a confident, wrong AI agent answer to bad context. Governed semantic ...
Some results have been hidden because they may be inaccessible to you
Show inaccessible results