A new benchmark, Mage-Bench, tests the strategic reasoning of top AI models by forcing them to play Magic: The Gathering, a game of immense complexity, hidden information, and long-term planning.
Research WireMany developers have a graveyard of unfinished side projects. A new perspective reveals how AI can act as the perfect co-pilot, crushing burnout and blockers.
Builder LabResearchers have unveiled a new AI model that infers the underlying reasons for your online activity to deliver hyper-personalized news recommendations across domains.
Research WireResearchers have unveiled the first-ever scaling laws for discrete diffusion language models, discovering a novel training method that makes them 12% more efficient.
Research WireA new paper reveals Boundary Point Jailbreaking (BPJ), a novel attack that can bypass even the strongest, human-tested LLM safety filters. Unlike prior methods, BPJ works on black-box models without needing internal access, posing a significant real-world threat.
AIBW Security DeskThe race for larger LLM context windows has a hidden cost. A new benchmark study reveals that as context grows, models struggle to focus, hindering personalization and increasing privacy risks. Is bigger always better?
Research WireThe next blockbuster drug may be hidden in a Chinese patent or a Russian journal. A new AI agent system called 'Hunt Globally' is designed to find it, breaking barriers.
Research WireVision-language models excel at understanding color photos but fail on thermal imagery. A new benchmark, ThermEval, aims to fix this critical gap for real-world tasks.
Research WireResearchers have uncovered why LLMs organize concepts like months into circles and years into lines. This geometric structure isn't random; it's a direct result of 'translation symmetries' hidden within the statistics of human language itself.
Research WireA new paper reveals a breakthrough in AI text style transfer. By translating text to another language and back, researchers create a 'neutral' version, solving the critical lack of training data. This enables LLMs to master styles like formality or politeness with remarkable efficiency.
Research WireA groundbreaking paper reveals the Distributed Quantum Gaussian Process (DQGP), a model merging quantum computing's power with AI to revolutionize multi-agent systems.
Research WireA newly digitized 18-year diary from a U.S. Forest Service worker offers a searchable window into American history from 1927 to 1945. The project, showcased on Hacker News, leverages modern tech to make this unique primary source accessible to all.
Builder Lab