4DAnyone: an open framework turns single-camera video into a 4D model of a person

24 August 2026
4DAnyone Create Anyone in 4D from a Casual Monocular Video

4DAnyone: an open framework turns single-camera video into a 4D model of a person

Researchers from Zhejiang University, Robbyant, Ant Group and HKUST introduced 4DAnyone, a framework that turns a video of a person shot on a single camera into a 4D model of…

Microsoft releases Agent Lightning v1.0 for training agents inside their own harness

20 August 2026
agent-lightning v1.0

Microsoft releases Agent Lightning v1.0 for training agents inside their own harness

Researchers from Microsoft, together with colleagues from Fudan University, Zhejiang University, and the University of Edinburgh, released Agent Lightning v1.0, a framework for reinforcement learning on LLM agents that fits…

JoyAI-Video-Edit: a 16B model brings real-time video editing to 30 FPS at 720p on a single B200

11 August 2026

JoyAI-Video-Edit: a 16B model brings real-time video editing to 30 FPS at 720p on a single B200

The Joy Future Academy team published JoyAI-Video-Edit, an open autoregressive diffusion model with 16 billion parameters that performs instruction-guided video editing on a live stream, with no access to future…

Kimi K3 Review: Moonshot AI’s Architecture, Benchmarks, Open Weights, and Fable 5 Comparison

3 August 2026
kimi k3

Kimi K3 Review: Moonshot AI’s Architecture, Benchmarks, Open Weights, and Fable 5 Comparison

Moonshot AI has released Kimi K3, the first open model in the 3-trillion-parameter class, with a 1-million-token context window. It is already available in the Kimi chatbot and through the…

TurboVLA: Robot Gets Commands Several Times Faster at a 97% Success Rate Thanks to Swapping the LLM for BERT

3 August 2026
TurboVLA model robot munipulation

TurboVLA: Robot Gets Commands Several Times Faster at a 97% Success Rate Thanks to Swapping the LLM for BERT

Researchers from Huazhong University of Science and Technology and Huawei have released TurboVLA, a compact vision-language-action model that produces robot actions in 31.2 ms on a consumer-grade RTX 4090 and…

Bonsai 27B: 1-Bit Weights Put a 27B-Parameter Model on a Smartphone for the First Time

15 July 2026
Bonsia 27B 1-bit

Bonsai 27B: 1-Bit Weights Put a 27B-Parameter Model on a Smartphone for the First Time

PrismML, a startup founded by Caltech researchers, has announced Bonsai 27B — binary and ternary versions of the Qwen3.6-27B model that retain 90–95% of the original model’s quality while compressing…

MIRA: A World Model Fully Simulates Rocket League Without Requiring You to Install the Game Itself

8 July 2026
MIRA world model rocket league simulation AI

MIRA: A World Model Fully Simulates Rocket League Without Requiring You to Install the Game Itself

Teams from General Intuition, Kyutai, and Epic Games introduced MIRA — a world model that fully simulates the Rocket League game environment for four players at once and draws each…

Claude Sonnet 5: A Strong Agentic Upgrade, but No Clear Opus Replacement

1 July 2026
claude sonnet 5 new

Claude Sonnet 5: A Strong Agentic Upgrade, but No Clear Opus Replacement

Anthropic has introduced Claude Sonnet 5, a new model in the Claude family that is also available to users on the free tier. It is designed for agentic tasks, programming,…

LFM2.5-230M: An Ultra-Compact Model Runs on a Raspberry Pi and Almost Any Modern Phone

29 June 2026
liquid AI

LFM2.5-230M: An Ultra-Compact Model Runs on a Raspberry Pi and Almost Any Modern Phone

Liquid AI released LFM2.5-230M — one of the smallest language models out there today, at just 230 million parameters. It’s compact enough to run on a small device without trouble:…

DreamX-World-5B: An Open-Source World Model with Camera Control, Text-Based Control, and Location Memory

17 June 2026
DreamX-World-5B модель

DreamX-World-5B: An Open-Source World Model with Camera Control, Text-Based Control, and Location Memory

The AMAP-ML team has published DreamX-World 1.0, an interactive generative world model that turns text or an image into a controllable video with precise camera control, memory of previously visited…

VibeThinker: 3B model reasons and codes at the level of flagship models

16 June 2026
https://neurohive.io/ru/ii-v-marketinge/pochemu-socseti-blokirujut-67-multiakkaunterov-v-pervye-3-dnya-analiz-500-otchetov/

VibeThinker: 3B model reasons and codes at the level of flagship models

Sina Weibo AI published VibeThinker-3B — a compact language model with just 3 billion parameters that matches flagship models DeepSeek V3.2 (671B), GLM-5 (744B), and Gemini 3 Pro on verifiable…

ESM Cambrian: protein language model outperformed Google’s AlphaFold3 and built the largest atlas of the protein world

4 June 2026
esm model

ESM Cambrian: protein language model outperformed Google’s AlphaFold3 and built the largest atlas of the protein world

A team of researchers from Biohub published ESM Cambrian (ESMC) — a language model for protein structure prediction and design that outperformed AlphaFold3 by Google on structure prediction accuracy, designed…

LLaVA-OneVision-2: Multimodal Model Analyzes Compressed Video Stream Through a Codec Instead of Frame Sampling

28 May 2026
LLaVA-OneVision-2

LLaVA-OneVision-2: Multimodal Model Analyzes Compressed Video Stream Through a Codec Instead of Frame Sampling

Researchers from Glint Lab, AIM for Health Lab, and MVP Lab published LLaVA-OneVision-2 (LLaVA-OV-2) — a next-generation multimodal model that rethinks how a neural network “watches” video. Instead of slicing…

SenseNova-U1: NEO-unify multimodal architecture works directly with pixels without VAE

14 May 2026
SenseNova-U1 Unifying Multimodal model

SenseNova-U1: NEO-unify multimodal architecture works directly with pixels without VAE

SenseNova introduced a new multimodal architecture, SenseNova-U1, which combines image understanding, generation, and editing inside a single transformer without a separate visual encoder or variational autoencoder. This approach removes the…

OpenSeeker-v2: Best-in-Class Deep Research Agent Built by an Academic Team on Just 10,600 Samples

7 May 2026
OpenSeeker-v2

OpenSeeker-v2: Best-in-Class Deep Research Agent Built by an Academic Team on Just 10,600 Samples

Researchers from Shanghai Jiao Tong University have proven that building a best-in-class deep research agent doesn’t require hundreds of billions of pre-training tokens or expensive reinforcement learning. Just 10,600 carefully…

OpenGame: AI Agent Generates Full Browser Games from Text Description

22 April 2026
gameengine

OpenGame: AI Agent Generates Full Browser Games from Text Description

A team of researchers from CUHK MMLab published OpenGame — the first agentic framework for creating browser-based 2D games from natural language descriptions. The project is fully open: the framework…

ClawGUI: the first open-source end-to-end framework for GUI agents — from training to real device

15 April 2026

ClawGUI: the first open-source end-to-end framework for GUI agents — from training to real device

Researchers from Zhejiang University have published ClawGUI — a fully open-source framework for building GUI agents that control applications through their visual interface, just like a human would: taps, swipes,…

ClawBench: The Best AI Agent Completed Only 33% of Real Everyday Online Tasks

13 April 2026

ClawBench: The Best AI Agent Completed Only 33% of Real Everyday Online Tasks

ClawBench — a benchmark testing whether AI agents can complete real everyday online tasks: booking a flight, applying for a job, placing an order. Results showed that even the strongest…

InCoder-32B-Thinking: Open-Source Code Generation Model for Microcontrollers, GPU Kernel Optimization, and RTL Design

7 April 2026
Overview of InCoder-32B-Thinking

InCoder-32B-Thinking: Open-Source Code Generation Model for Microcontrollers, GPU Kernel Optimization, and RTL Design

A research team from Beihang University, Shanghai Jiao Tong University, the University of Manchester, and IQuest Research has published InCoder-32B-Thinking — a language model with an extended chain-of-thought reasoning for…

Trinity-Large-Thinking 400B: an open model matching Claude Opus-4.6 on agentic benchmarks at 28x lower price

3 April 2026
Trinity AI models foundation

Trinity-Large-Thinking 400B: an open model matching Claude Opus-4.6 on agentic benchmarks at 28x lower price

Arcee AI has released Trinity-Large-Thinking — an open-weight reasoning model for complex multi-turn agentic tasks. On PinchBench — a comprehensive benchmark for AI agents — it ranks second among all…

PixelSmile: Open Model for Facial Expression Editing with Smooth Intensity Control

31 March 2026
PixelSmile

PixelSmile: Open Model for Facial Expression Editing with Smooth Intensity Control

Researchers from Fudan University and StepFun have published PixelSmile — a diffusion model for precise facial expression editing in portraits and anime images. Instead of training on discrete labels like…