Updated
#Embodied AI
Updated
Aug 24
hardmaruAI score23
Aug 21
Jim FanAI score59 NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method
AINVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.
Aug 20
hardmaruAI score13 Winning at all costs
Aug 19
Jazzyear · InsightsAI score36 Zhang Yijia Forecasts 2026 AI Trends: Capital Surge and Next-Generation Intelligence Paradigms
AIZhang Yijia, founder and CEO of Chinese tech think tank 甲子光年, presented a 2026 AI trend report at a Beijing investment conference, arguing China's venture capital has entered a new cycle as financing, investment, and exits rebounded in H1 2026. He said AI absorbed over 70% of global venture investment in H1 2026, with OpenAI and Anthropic together raising $217 billion, and that global AI-related capex is expected to exceed $1 trillion in 2026.
GeneralistAI score42 To us, GEN-1.5 represents a new frontier of generality — one that challenges our own understanding of how these models behave when…
AI…pretrained at a scale of physical interaction data few thought possible without shortcuts. We do not yet see where this asymptotes. Read more in the full blog:
GeneralistAI score20 Or when fine-tuned to place a block into a bowl, it can clear obstacles (like a piece of paper covering the bowl) to complete the task…
AI…despite that not being in the demonstrations.
GeneralistAI score31 Fine-tuned (or prompted) behaviors generalize beyond their demonstrations, and can improvise fundamentally different manipulation…
AI…strategies to achieve the same goal. For example, after fine-tuning to use a brush to sweep a block into a bowl, it could use other tools like a dustpan to accomplish the same task with a very different strategy.
GeneralistAI score34 For few-shot learning, it can adapt to new physical tasks in as few as 1 - 10 gradient steps on 1 - 5 minutes of data (~10 - 50…
AI…demonstrations). In practice, this can be described as test-time training in a low-data regime. We did not tune this procedure or sweep hyperparameters; these results come largely out of the box.
GeneralistAI score46 Generalist model learns physical tasks from one or few demonstrations
AIGeneralist's model reached 59% average success on 10 diverse physical tasks with one-shot prompting straight from pretraining. With few-shot learning, using 10 gradient steps on 5 minutes of data per task, performance rose to 83%. The post calls it the first model it knows of that learns a wide range of dexterous closed-loop physical tasks from one or few demonstrations.
GeneralistAI score34 In some cases, in-context learning with GEN-1.5 transfers across the embodiment gap entirely: a human demonstrates a task with their own…
AI…hands, observable through the robot’s cameras, and the robot can reproduce it immediately afterward.
GeneralistAI score38 In-context learning also crosses the sim-to-real gap, zero-shot.
AIPrompts can be formed entirely from simulated experience (e.g., from a scripted policy, an RL agent, or a human teleoperating a simulated robot) and be used to produce behaviors on a real robot. The model was not trained on the task in either the simulator or the real world.
Aug 17
Rowan CheungAI score34 Keio and MIT Media Lab unveil a floating helium-filled companion robot
AIResearchers from Keio University and the MIT Media Lab built a soft, helium-filled robot with gentle flapping fins and no face, rotors, or pinch points. In demos it served as an alarm clock, study buddy, movement reminder, and dance partner, communicating through movement. The team designed it to avoid the uncanny valley and be safe to touch, betting people will welcome it into their personal space.
Jul 28
Jazzyear · InsightsAI score28 GDPS 2026 launches global calls for scenarios, teams, and experts for embodied AI contest
AIThe 2026 Global Developers Pioneer Conference (GDPS) and International Embodied Intelligence Skills Competition has opened global calls for scenario challenges, teams, and experts ahead of its October 23–25 event at Shanghai Zhangjiang Science Hall.
Tri DaoAI score42 Putting LLM brain on robots -> 4x SOTA with no extra training.
AII’ve been very surprised by how well this works. The time for agents running on robots is coming soon
Fei-Fei LiAI score34 When SceniX joined World Labs, we said spatial intelligence was never only about perceiving and generating virtual and physical worlds, but…
AI…also interacting with them. Today, we’re sharing early results from that vision: building worlds that train robots. 🌎🤖↓
Jul 27
Rowan CheungAI score40 Weave Robotics $8,000 robot that folds laundry and makes your bed just went up for preorder, deliveries slated for this fall.
AIRobot butler season is coming.
Air Street PressAI score72 Black Forest Labs releases FLUX 3, extended to video and robot control
AIBlack Forest Labs released FLUX 3, a multimodal model trained on images, video, and audio, and mimic built FLUX-mimic on its video backbone to control robots. In a soft-body kitting task, mimic reports a 95% success rate without single-task fine-tuning, compared with 55% for an adapted π0.5 model. FLUX 3 Video is in early access, with action prediction offered to selected partners and an open-weight backbone planned.
Meta AI BlogAI score36 Meta's DINOv3 and SAM Power Edge-Based Assistive Robotics at Pittsburgh
AIThe University of Pittsburgh's RAMMP team is integrating Meta's DINOv3 and SAM models into on-device assistive robotics to detect door buttons, cups, and curbs for navigation assistance. The models run on compact, battery-powered hardware, with optimizations such as reduced memory footprint and lower precision, enabling real-time perception without network connectivity. RAMMP's perception system pairs SAM-based auto-labeling with an RF-DETR detector fine-tuned on DINOv2 embeddings, and the team is now testing voice and touch input for object selection.
Jul 21
Fei-Fei LiAI score33 The world is not just made of words, and spatial intelligence was never just about perceiving and generating worlds.
AIIt's about interacting with them. Today, SceniX is joining World Labs. 🌎🤖👇
Air Street PressAI score67 DeepMind's Raia Hadsell argues AI should move beyond language to world models and robotics
AIAt RAAIS, DeepMind VP of Research Raia Hadsell argued that the field focuses too much on language and should apply large-model training to worlds, robots, biology, and weather. The article cites DeepMind's DiffusionGemma, a 26-billion-parameter open text model that generates blocks by denoising rather than one token at a time, and the Genie-3 world model, which runs in real time for several minutes. It also describes world models as a source of synthetic training data for robots.
Jul 20
Jul 17
Jim FanAI score13 My morning ritual: coffee, then watch our robot assemble stuff.
AIThe model isn't fast, but it measures every grasp, obsesses over every alignment, and handles every piece like an heirloom. Watching is therapeutic, even meditative. Simple pleasure from the execution of a task well done. Uncut, no speedup, end-to-end policy in one go.
Jul 15
Fei-Fei LiAI score60 RoboTTT scales robot policy context to 8,000 timesteps using test-time training
AIStanford SVL and NVIDIA Robotics introduced RoboTTT, which uses test-time training to give robot policies up to 8,000 timesteps of context at constant inference cost. The source reports that 8K-context pretraining beats 1K by 62%, and that performance keeps improving from 128 to 8K timesteps with no sign of saturation. The authors also describe one-shot imitation from human video and in-episode error recovery.
Jim FanAI score62 RoboTTT scales robot policy context to 8,000 timesteps with constant inference cost
AIJim Fan introduced RoboTTT, a robot model that uses test-time training to compress history into a tiny inner model updated at each sensor reading. The post reports closed-loop performance rising steadily from 128 to 8K timesteps, and 8K-context pretraining beating 1K by 62%. It also claims one-shot in-context learning from human video and mid-episode error recovery, with learning continuing after deployment.
Jul 9
BAAIAI score47 BAAI releases Orca, a multimodal latent world model using Next-State-Prediction
AIBAAI introduced RoboBrain Orca, which learns world latent representations from multimodal data and models world-state transitions through Next-State-Prediction instead of only next-token, next-frame, or next-action prediction.
Jul 1
Jim FanAI score51 Jim Fan introduces ASPIRE, a self-evolving robot skills library for continual learning
AIJim Fan announces ASPIRE, a system where coding agents use multimodal sensory traces from simulation and real robots to run evolutionary search over control programs and add the results to a growing skills library. The post claims up to a roughly 10x reduction in transfer learning tokens for sim-to-real and single-arm to bimanual transfer, and says the full stack will be open-sourced.
Jun 30
Jim FanAI score60 ASPIRE lets robots build an evolving skills library that transfers across tasks
AIJim Fan introduces ASPIRE, a system in which coding agents observe multimodal sensory traces and run evolutionary search over control programs to distill skills into a growing library. The post says ASPIRE shares know-how rather than pixels or weights across the sim-to-real gap, reducing transfer learning tokens by up to about 10x. The author also says the full stack will be open-sourced and provides a gallery of 150+ tasks and 90+ skills.
Jun 20
BAAIAI score14 🔬 Immersive AI Research Experience Zone The 8th BAAI Conference unveiled an immersive zone featuring SoulAgent, embodied intelligence…
AI…FlagOS, and life-science AI for Science. Hands-on demos included humanoid robot task demonstrations, multi-chip system experiences, and AI-assisted cardiac diagnosis, neuroscience, and drug discovery—making real-world AI tangible. #BAAIConference #AI
Jun 17
Jim FanAI score64 ENPIRE lets Codex agents run autonomous research on a robot fleet
AINVIDIA GEAR's ENPIRE gives eight Codex agents a fleet of robots, GPUs, and a token budget to solve physical tasks with minimal human oversight. The author reports tasks such as tying zip-ties, organizing fine pins, and installing GPUs, and a faster time-to-solution with eight parallel robots than with fewer. Safety uses a kinematic limit that resets a robot leaving its envelope, a torque-limited gripper, and a frozen reward function classifier. The team says everything will be open-sourced.
Jun 16
Jim FanAI score62 Jim Fan's ENPIRE lets Codex agents run autonomous research on robot fleets
AIJim Fan introduces ENPIRE, which gives eight Codex agents a fleet of robots, GPUs, and a token budget to solve physical tasks autonomously. The post reports that the system can tie zip-ties, organize fine pins, and install GPUs, and that eight robots exploring in parallel improve faster than fewer. The team plans to open-source everything.
Jun 5
Jim FanAI score34 NitroGen just won CVPR Best Paper Honorable Mention!!
AIWe are making strides towards general-purpose embodied agents that master not only the real world physics, but also all possible physics across a multiverse of simulations. It’s been 4 years since MineDojo, our first embodied agent in Minecraft, won NeurIPS Best Paper. Congrats to everyone on the team!!
BAAIAI score20 BAAI's 8th Conference set for June 12–13 in Beijing
AIThe 8th BAAI Conference will be held June 12–13 at the Zhongguancun Innovation Center in Beijing, with Turing Award laureates and China's large-model leaders attending. Core focuses include world models and agents, plus two new flagship sessions on AI-native education and the token economy. The event will feature 25 forums, over 200 speeches, and a first-ever on-site AI agent conference companion for real-time listening and summarization.
Apr 17
NVIDIA AI DeveloperAI score23 🙌 Meet the developers who won the Cosmos Cookoff — and discover how they used Cosmos Reason 2 to build breakthrough projects: 🤖from…
AI…disaster-response drones to explainable visual AI and intelligent security systems. 🎥 Watch the showcase → 📖 Read the full recap →
Apr 1
Jim FanAI score62 CaP-X open-sources agentic robotics toolkit, benchmark, and RL setup
AIJim Fan announced the open-source release of CaP-X, an agentic robotics framework in which LLM-driven agents control robot arms and humanoids through perception and actuation APIs. The release includes CaP-Gym with 187 manipulation tasks across RoboSuite, LIBERO-PRO, and BEHAVIOR, and CaP-Bench, which evaluates 12 frontier LLMs and VLMs across 8 tiers. The post also reports that a 7B open-source model rose from 20% to 72% success after 50 RL training iterations, with synthesized programs transferring to real robots.
Mar 23
Jim FanAI score40 Teleop is so 2025. Ever since we unveiled EgoScale and the dexterity scaling law, it's been clear to us and the ecosystem that behavior…
AI…cloning directly from humans is the way to break the curse of teleop. 2026 is all about scaling robot learning without robots.
Mar 17
BAAIAI score46 BAAI unveils RoboBrain-Dex, dexterous manipulation trained on human egocentric data
AIBAAI has released RoboBrain-Dex, a dexterous manipulation model for embodied intelligence trained on large-scale, diverse human egocentric data rather than massive robot teleoperation datasets. The company says this approach yields strong generalization, marking a shift from small-data, weakly generalizing methods toward big-data robotic manipulation. The code has been open-sourced on GitHub.
Feb 25
Jim FanAI score75 EgoScale trains a 22-DoF humanoid mostly on 20,000 hours of human video
AIResearchers trained a humanoid with 22-DoF dexterous hands mainly on over 20,000 hours of egocentric human video, with no robot in the loop, to perform tasks such as assembling model cars and folding shirts. They report a log-linear scaling law (R² = 0.998) between human video volume and action prediction loss, and state that this loss predicts real-robot success rate. The recipe, called EgoScale, pre-trains GR00T N1.5 on the video, adds only 4 hours of robot play data, and reports a 54% gain over training from scratch across five dexterous tasks.