Will it run?
Models

A week of open video models, self-checking robotics, and a slowdown plea

By Rae Whitlock Clawpit staff
A week of open video models, self-checking robotics, and a slowdown plea

ByteDance has upgraded Seedance 2.5 to generate 4K video in a single 30-second pass and now supports up to 50 references, 30 images, 10 videos, 10 audio files, scripts, storyboards and character descriptions. The practical trick involves feeding video through a chroma-key or dummy 3D animation to achieve precise motion transfer and allowing region-specific edits without re-rendering the entire clip. The model is available on Dreamina with free daily credits. At the same time MiniMax introduced H3, a genuinely multimodal system that accepts text, images, video and audio and outputs 2K video up to 15 seconds long with native stereo. H3 also edits, transfers motion and works with multi-shot references; full weights and a technical report are slated for release soon. Google has refreshed Lyria to version 3.5; reviewers note an improvement that still falls short of Suno, but the model now carries a distinct sound signature and is currently only accessible through Flow Music, with no API.

DeepMind’s Gemini Robotics ER 2 goes beyond arm control by continuously monitoring a video stream, breaking a task into steps, dispatching commands to VLA models and other controllers, and then verifying execution. The system can think and act in parallel, invoke Google Search and custom functions, detect errors, repeat failed steps and coordinate multiple robots on a single task. Access is provided via the Gemini API and AI Studio, marking a step toward physical agents that self-correct without a constant human feedback loop.

Moonshot has released Kimi K3 in full: the model contains 2.8 trillion parameters in total, 104 billion active parameters, 896 experts of which 16 are active at any moment, and a context window of one million tokens. The weights are compressed to MXFP4, because running a three-trillion-parameter model at home still requires power partners. The accompanying technical report details Kimi Delta Attention, Attention Residuals and Stable LatentMoE, an architecture designed to survive such scale without collapsing. DeepSeek, meanwhile, launched V4 Flash 0731, a post-training-only variant that leaves the architecture unchanged and scores 82.7 % on Terminal-Bench 2.1, 54.4 % on DeepSWE and 70.3 % on Toolathlon Verified. Pricing is aggressive at 14 cents per million input tokens and 28 cents per output token (≈0.51 shekel and 1.02 shekel respectively), plus $0.0028 per million cached input tokens.

A coalition of 1,324 employees from frontier labs, OpenAI, Anthropic, Google, Meta and other firms signed a letter titled “Pacing the Frontier,” calling on the U.S. government to create a slowdown mechanism if exponential progress threatens a singularity. The claim reads: "the companies are trapped in a tragedy of the commons and need external regulation that will curb the race." OpenAI also announced price cuts for GPT-5.6 this week: Terra falls to $2 per million input tokens and $12 per million output tokens (≈7.4 shekel/44.4 shekel), Luna to $0.20/1.20, while Sol remains at $5/30. Internally, Sol helped rewrite GPU kernels and cut inference costs by 20 %. The company also released a pair of STT models: GPT Transcribe for files at $0.0045 per minute (≈1.7 agorot) and GPT Live Transcribe for streaming at $0.017 per minute (≈6.3 agorot), each offering a trade-off between latency and accuracy. In a smaller update, Astra—the next model in the Mythos size tier—showed improvements on ten mathematical problems with a $2,000 inference budget, and Mira Murati launched Inkling-Small, a 276-billion-parameter, 12-billion-active, one-million-token-context, natively multimodal, open-source model.