Text to Video News and AI Updates
Follow Text to Video developments across AI companies, labs, and open-source projects.
Latest Text to Video news, research, benchmarks, product releases, and industry adoption updates.
Latest Articles
- LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models — LPA-CWM adds a lightweight 3M-parameter adjudicator to counterfactual world models for video motion tracking, predicting response reliability rather than using uniform aggregation—improving tracking accuracy by 60% on DAVIS and 29% on Kinetics without modifying the underlying video model.
- Dynamic Learning Solutions: A System for Personalized Educational Video Generation — Researchers built a pipeline that converts PDF textbooks into animated video explanations by chaining RAG retrieval, Stable Diffusion image generation, DynamiCrafter animation, and Google TTS — demonstrated on NCERT educational materials.
- BEACON: Behavior and Appearance Control for Subject-Specific Video Generation — BEACON generates person-specific videos that preserve both visual identity and characteristic facial dynamics by disentangling the two into separate conditioning signals — updating only ~1% of the Wan video diffusion model trained on 2,000 pairs, with improved expressivity over state-of-the-art methods.
- New York Seizes a Dozen Celebrity Deepfake Websites — The Manhattan DA's Office seized 12 nonconsensual deepfake pornography websites targeting ~1,200 victims — the largest such enforcement action to date — using the US Take It Down Act, which now gives law enforcement authority to seize domains hosting synthetic sexual imagery.
- MindTopo: Can Foundation Models Reason in Topological Space? — MindTopo — 11,030 instances across continuity, separation, order, enclosure, and knot tasks — finds all 14 MLLMs tested fall far below human performance on topological planning, with fine-tuning improving reasoning more than planning and video generation failing to preserve topology across transitions.
- From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models — VWG-Bench exposes a stark gap in video generation: models score well on visual quality but consistently fail logic-heavy and rule-constrained tasks across 9 reasoning dimensions. Vid-PRE, a RL-trained prompt rewriter, closes much of this gap without touching model weights.
- AgenticGen: Reward-Guided Agentic Video Generation for Advertising — TikTok's AgenticGen uses DPO followed by GRPO to optimize advertising video generation against live business metrics, achieving +2.72% CTR, +2.63% CVR, and +9.61% advertising value over SFT baseline in production A/B tests.
- San Francisco Orders Meta to Stop ‘Allowing’ AI Child Abuse Ads — San Francisco's city attorney sent Meta a cease-and-desist after the company ran 350+ paid ads featuring AI-generated child sexual abuse imagery on Facebook and Instagram — including depictions of confirmed real children — with 250+ ads continuing to run even after Wired reported the first batch in August.
- Beyond Coherence: Benchmarking Professional Editing-Technique Execution in Multi-Shot Audio-Video Generation — CutCraft is the first benchmark testing whether AI video generators actually execute professional editing techniques — J-cuts, L-cuts, shot transitions — finding that current SOTA models produce plausible videos but consistently fail to follow editorial instructions.
- AV-SafetyBench: A Safety Benchmark for Text-to-Audio-Video Generation — AV-SafetyBench is the first safety benchmark for text-to-audio-video generation, covering 5,200 prompts across 13 categories — finding that 41.6-48.3% of unsafe outputs from T2AV models are missed entirely by video-only evaluation, and 87.5% of cross-modal harms emerge only from joint audio-video interpretation.