Skip to contentSkip to stories

Updated

#Embodied AI

Oct 8

Oct 8Thu

Oct 7

Oct 7Wed

Oct 6

Oct 6Tue

Sep 30

Sep 30Wed

Sep 22

Sep 22Tue
  1. Black Forest Labs · new models on Hugging FaceAI score62

    Black Forest Labs releases FLUX 3 Action, a 7B open-weights robot world action model

    AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.

    Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.

  2. Black Forest Labs · new models on Hugging FaceAI score58

    Black Forest Labs releases open-weights FLUX 3 Action SO-101 robot policy

    AIBlack Forest Labs has published FLUX 3 Action SO-101 on Hugging Face as an open-weights 7B world action model. It takes two camera frames, the robot state, and a text instruction, then returns the next 42 actions with predicted video frames, with 32 executed at 30 Hz before replanning. The card also provides a rank-32 LoRA fine-tuning recipe for user datasets and states that the application must enforce joint velocity, force, and workspace limits.

  3. Black Forest Labs · new models on Hugging FaceAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

Sep 15

Sep 15Tue

Sep 14

Sep 14Mon
  1. NVIDIA · new models on Hugging FaceAI score36

    NVIDIA's FoundationPose estimates 6-DoF object pose without fine-tuning given a CAD model

    AINVIDIA released FoundationPose, a transformer-based model for 6-DoF object pose estimation and tracking that works on novel objects at test time without fine-tuning, given a CAD model. It takes RGB and depth images, a 2D bounding box, a CAD model, and camera intrinsics as inputs, and is licensed under the NVIDIA Open Model License for commercial use. The model is trained on synthetic data from Objaverse and Google Scanned Objects, with evaluation on LINEMOD and YCB-Video.

Sep 2

Sep 2Wed
  1. NVIDIA · new models on Hugging FaceAI score36

    NVIDIA Releases EgoHand-1.0 Model for Single-Image 3D Hand Pose Estimation

    AINVIDIA released EgoHand-1.0, a 883.5M-parameter DINOv3-based transformer that predicts SOMA hand pose, MHR shape coefficients, and camera translation from a single 256×256 hand crop. The model is evaluated on the HOT3D egocentric benchmark and is intended for research and demonstration rather than production use. Its outputs can supply hand trajectories for training robotic manipulation policies, and it runs on NVIDIA Ampere GPUs under Linux with PyTorch.

Sep 1

Sep 1Tue

Aug 26

Aug 26Wed

Aug 24

Aug 24Mon

Aug 19

Aug 19Wed

Jul 27

Jul 27Mon
  1. Air Street PressAI score72

    Black Forest Labs releases FLUX 3, extended to video and robot control

    AIBlack Forest Labs released FLUX 3, a multimodal model trained on images, video, and audio, and mimic built FLUX-mimic on its video backbone to control robots. In a soft-body kitting task, mimic reports a 95% success rate without single-task fine-tuning, compared with 55% for an adapted π0.5 model. FLUX 3 Video is in early access, with action prediction offered to selected partners and an open-weight backbone planned.

Feb 24

Feb 24Tue
  1. Jim FanAI score62

    NVIDIA's SONIC trains a 42M transformer to control a humanoid robot

    AINVIDIA researchers trained SONIC, a 42M-parameter transformer, to control a humanoid robot's whole body using motion tracking on over 100M mocap frames. After three days of training in simulation, the policy transferred zero-shot to the real G1 robot and reported a 100% success rate across 50 real-world motion sequences. One policy supports VR teleoperation, webcam human video, text prompts, music, and GR00T N1.5 VLA integration with 95% success on mobile tasks, and the code and checkpoints are open-sourced.

Jan 26

Jan 26Mon
  1. BAAIAI score40

    BAAI RoboBrain 2.5 targets robot spatial and temporal reasoning gaps

    AIBAAI released RoboBrain 2.5, an embodied AI model that turns 2D scene understanding into actionable 3D trajectories and provides dense temporal value estimates for real-time progress feedback on long-horizon tasks. The post says it achieves SOTA across multiple spatial and temporal reasoning benchmarks, though it names no specific scores. Project page, paper, GitHub code, and model weights are linked.