Large Language Models Benchmarks

Measuring What Matters in Large Language Model Performance

As large language models (LLMs) gain momentum worldwide, there’s a growing need for reliable ways to measure their performance. Benchmarks that evaluate LLM outputs allow developers to track ...

If you code Android apps with AI, Google’s new benchmark makes it easier to pick the right model

For Android app developers relying on AI to code, picking the right model can be tricky. Not all models are built the same, and many are not specifically trained for Android development workflows. To ...

IFLScience

"Humanity's Last Exam" Reveals How Accurate AI Actually Is. Chatbots Might Want To Look Away Now.

In updated tests published to the Humanity's Last Exam website, Gemini's 3.1 Pro model achieved 45.9 percent accuracy, with a ...

Qwen 3.5 35B vs Sonnet 4.5 : Benchmarks vs Reality Results Across Three Tasks

The rivalry between Qwen 3.5 and Sonnet 4.5 highlights the shifting priorities in large language model development. Qwen 3.5, ...

Why ‘winning’ the AI race is so hard to define

AI development is often framed as a race among countries, companies and academic researchers. But figuring out who’s actually ...

14h

Hey ChatGPT, write me a fictional paper: these LLMs are willing to commit academic fraud

All major large language models (LLMs) can be used to either commit academic fraud or facilitate junk science, a test of 13 ...

Model Show: Coding, OCR, and Chinese New Year

February brought new coding models, and vision-language models impress with OCR. Open Responses aims to establish itself as a ...

Crypto Briefing

Anthropic launches AI exposure index to assess which white-collar jobs face automation risk

Anthropic's new AI Exposure Index ranks computer programmers as the most vulnerable to LLM automation, with 75% of tasks automatable and early-career hiring slowing.

1don MSN

Stop Guessing: Google Now Ranks the Best AI for Android Coding

The post Stop Guessing: Google Now Ranks the Best AI for Android Coding appeared first on Android Headlines.

HealthTech Magazine

OpenAI, HealthBench, Claude, and HIPAA Compliance: What Healthcare IT Needs to Know

There’s been a lot of movement in healthcare among companies behind major large language models. Here’s a look at how the ...

14h

Bond Investors May Be Underpricing Cyber Risk As AI Speeds Up Attacks

Exploit timelines have collapsed and AI is compressing them further. A growing body of research suggests credit and loan ...

OpenAI launches GPT-5.4 with computer vision, tool use enhancements

OpenAI Group PBC today launched a new large language model that it says is more adept at automating work tasks than its earlier algorithms. GPT-5.4 is available in ChatGPT, the Codex programming tool ...

Some results have been hidden because they may be inaccessible to you

Show inaccessible results