AI News

Latest AI engineering news, updated daily.

In-depth tutorials and guides. Go to Blog →

Ai Engineering

CyberSecQwen-4B Defeats Cisco 8B on CTI-MCQ Benchmark

Team athena19 fine-tuned a 4-billion parameter model on a single AMD MI300X GPU that outperforms Cisco's 8B model for defensive cyber threat intelligence.

Fine Tuning · Cyber Threat Intelligence · Model Optimization

Ai Engineering

EMO Pretraining Decouples Mixture-of-Experts Subsets

AI2 and UC Berkeley researchers introduced EMO, a pretraining constraint that groups MoE experts by semantic domain to allow independent subnet deployment.

Mixture Of Experts · Pretraining Methodology · Model Optimization

Ai Agents

Perplexity Opens Personal Computer Mac Agent to Pro Subscribers

Perplexity's new macOS application transitions out of preview, bringing hybrid cloud-local agent workflows to all Pro and Max tier users.

Perplexity Ai · Macos Agent · Personal Computer Agent

Ai Agents

AlphaEvolve Agent Refines Core Algorithms via Gemini Ensemble

Google DeepMind detailed real-world deployments of its AlphaEvolve coding agent, showing measured gains in quantum simulation, genomics, and infrastructure.

Autonomous Agents · Deepmind Gemini · Algorithm Optimization

Ai Engineering

Roche Integrates PathAI Diagnostic Algorithms in $1.05B Deal

Roche has acquired Boston-based PathAI in a $1.05 billion transaction to embed AI-powered image analysis directly into its global oncology diagnostic platforms.

Generative Ai · Healthcare Technology · Oncology Diagnostics

Ai Coding

Windsurf IDE Adds Local SWE-check and Cloud Devin Review

Cognition AI has integrated Devin Review and a local SWE-check model into the Windsurf IDE to automate code verification and complex pull request reviews.

Windsurf Ide · Devin Review · Swe Check

Ai Agents

IBM Bob Agent Automates the SDLC With Multi-Model Routing

At Think 2026, IBM launched the Bob SDLC agent system, enterprise agent control planes, and detailed its $11 billion acquisition of Confluent.

Ibm Think 2026 · Sdlc Automation · Agentic Ai

Ai Engineering

vLLM V1 Migration: Fix Logprobs Before RL Corrections

ServiceNow's vLLM V1 migration shows why RL pipelines need backend logprob parity before objective-level corrections.

Llm Inference · Vllm Framework · Reinforcement Learning

Ai Engineering

GPT-5.5 Instant Cuts ChatGPT Hallucinations by 52.5%

OpenAI has replaced ChatGPT's default engine with GPT-5.5 Instant, a less verbose model featuring improved factuality, personalization, and memory sources.

Llm Optimization · Hallucination Reduction · Openai Updates

Ai Engineering

Private Evaluation Track Deters Open ASR Benchmaxxing

Hugging Face partnered with Appen and DataoceanAI to introduce a private evaluation track to the Open ASR Leaderboard, mitigating test-set contamination.

Automatic Speech Recognition · Model Evaluation · Benchmarking Integrity

Ai Engineering

GENE-26.5 Gives Hardware-Agnostic Robots Human-Scale Dexterity

The French robotics startup Genesis AI has released GENE-26.5, a hardware-agnostic foundation model paired with a custom human-scale robotic hand.

Robotics · Foundation Models · Hardware Agnostic

Ai Agents

Claude Managed Agents Add Background Dreaming and Subagents

Anthropic updated Claude Managed Agents with background memory consolidation, multiagent orchestration, and rubric-based output grading for complex workflows.

Multiagent Orchestration · Memory Consolidation · Autonomous Workflows

Ai Engineering

Steering Chemical Synthesis via LLM Evaluation in EPFL's Synthegy

EPFL researchers have developed Synthegy, a framework that uses large language models to evaluate and guide traditional computational chemistry algorithms.

Computational Chemistry · Llm Evaluation · Chemical Synthesis

Ai Engineering

Native iOS 27 Workloads Can Now Route to Claude and Gemini

Apple's Extensions framework for iOS 27 allows developers to integrate third-party AI models directly into native Siri and Writing Tools workflows.

Ios Development · Apple Intelligence · Model Integration

Prompt Engineering

Mindgard Uses Visible Thinking to Jailbreak Claude Sonnet 4.5

Security firm Mindgard bypassed Claude Sonnet 4.5 safety filters by using psychological pressure to manipulate the model's visible internal reasoning process.

Adversarial Attacks · Jailbreaking · Large Language Models

Ai Agents

$27M Funding Round Backs CopilotKit's App-Native Agent Stack

CopilotKit has raised $27 million to expand its generative UI framework and launch a self-hostable enterprise intelligence platform for app-native AI agents.

Venture Capital · App Native Agents · Generative Ui

Ai Agents

AWS Tackles Agent Drift With Bedrock AgentCore Optimization

AWS has introduced AgentCore Optimization in preview to automate prompt updates and A/B testing, alongside a new desktop AI assistant called Amazon Quick.

Aws Bedrock · Agent Drift · Prompt Optimization

Ai Engineering

Runpod Flash Removes Container Overhead for AI Inference

The open-source Flash Python SDK allows developers to convert local functions into auto-scaling serverless AI inference endpoints without Dockerfiles.

Serverless Inference · Cloud Infrastructure · Open Source Sdk

Ai Engineering

DeepSeek V4 Pro Trails GPT-5.5 by 8 Months in NIST Benchmarks

The Center for AI Standards and Innovation evaluated DeepSeek-V4-Pro, placing its capabilities eight months behind U.S. frontier models while matching GPT-5.

Deepseek V4 · Nist Benchmarks · Llm Evaluation

Ai Engineering

TPU v5p Inference Speeds Triple With DFlash Block-Diffusion

Google and UCSD researchers released DFlash, a block-diffusion speculative decoding method that achieves a 3.13x average inference speedup on TPU v5p hardware.

Llm Inference · Google Tpu · Speculative Decoding