Paril Jain, co-founder and CTO of The Bot Company, and Quan Vuong, co-founder of Physical Intelligence met on the Runway AI Summit stage to…
AI…talk real-world evals, deploying robots in the field and where robotics will be in three years.
Updated
Updated
AI…talk real-world evals, deploying robots in the field and where robotics will be in three years.
AIAgents expose errors in imagination, and repairing those errors enables further learning. We believe this open-ended learning is a critical step towards superintelligence.
AIThis recursive loop delivers up to 91% relative gains over the StarCraft world-model baseline, alongside stronger robot coordination in simulation. Learn more, read the PROWL-2 paper, and let us know any questions!
AI…model, Praxis-1. Learn more about Praxis-1 and request early access:
AIThe seven of us here do the work of 50," says Vignesh Anand, co-founder of @innate_bot. We spent time with the team to capture what building means to them.
AIOur next robot should have arms so that it can participate in the production of microducks 😅😅😅
AIMicrosoft Research Asia – Singapore, opened July 24, 2025 as Microsoft's first Southeast Asian research lab, reports progress after its first year. The lab's work spans next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development. Its healthcare collaborations on multimodal and agentic AI for clinical decision-making are being deployed through partnerships across Singapore's healthcare ecosystem.
AISakana AI and the University of Tokyo introduced SAIL, a method that generates robot trajectories with a VLM and refines them through simulator testing, VLM feedback, and Monte Carlo tree search. Across six simulated manipulation tasks, raising the search budget from one candidate to 45 increased the success rate of finding a working trajectory from 25% to 73%. The authors also tested the approach on a physical robot, though the post frames further transfer to real hardware as an open question.
AIAnd it connects to ChatGPT and Entire. Come meet Marvin, Stefano and us @wearedevs San Jose, booth 753. 🌠
AINew Microsoft Research findings show that moving AI inference beyond the robot can improve task success, boost efficiency, and support more advanced physical AI workloads.
AIMicrosoft Research reports that running physical AI inference on onboard GPUs can limit robot performance and battery life, while offloading inference to edge or cloud GPUs improved results in mobile manipulation tests. In its evaluation, smaller onboard GPUs slowed mapping and planning by up to 383% compared with an A100, and large onboard GPUs such as Jetson Thor drained robot batteries by up to 160%.
Why it matters: The study measures how offloading robot inference to edge or cloud GPUs changes task success, battery life, and model size, offering evidence for infrastructure design.
AILiveKit x @dimensionalos Robotics Happy Hour, #SFTechWeek Tue Oct 6 · 5:30pm · SF RSVP :
AINo download. Point your coding agent at it and build custom viewers fast. As an example, here's a cool UI implemented in ~600 lines of JS on top of the package ⤵️ (GitHub repo in reply) Kudos to @mishig25 and team LeRobot!
AIBlack Forest Labs released FLUX 3 Action, an open-weights 7B world action model that outputs robot joint commands from camera frames, robot state, and a text instruction. On the RoboLab-120 benchmark it reports 42.92% task success, ahead of Cosmos3-Nano-Policy at 36.8% and π0.5 at 28.0%. The model is fine-tuned on DROID, is distributed under the FLUX Kommunity License v.1.0, and runs in about 32 GB of GPU memory in bfloat16.
Why it matters: The model card gives a benchmark comparison, parameter counts, and an action contract, so readers can judge how it compares with existing robot policies.
AIBlack Forest Labs has published FLUX 3 Action SO-101 on Hugging Face as an open-weights 7B world action model. It takes two camera frames, the robot state, and a text instruction, then returns the next 42 actions with predicted video frames, with 32 executed at 30 Hz before replanning. The card also provides a rank-32 LoRA fine-tuning recipe for user datasets and states that the application must enforce joint velocity, force, and workspace limits.
AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.
Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.
AIRLinf, an open-source framework for embodied intelligence and AI agents, now supports Cosmos3 from fine-tuning through robot evaluation. With SGLang inference, it delivers 3.33x end-to-end evaluation throughput, batching inference for 128 parallel environments on 8 GPUs across 500 episodes of the full LIBERO-10 evaluation. RLinf also overlaps CPU simulation with GPU inference to reduce waiting between stages.
AI…with physical and virtual systems. We couldn't be more excited at the promise our models are showing.
AIIt can control robots, power humanoids, drive cars (on the roads of India!), train AIs, pilot drones, and even play video games. We can’t wait to see what intelligent systems it enables.
AINVIDIA released FoundationPose, a transformer-based model for 6-DoF object pose estimation and tracking that works on novel objects at test time without fine-tuning, given a CAD model. It takes RGB and depth images, a 2D bounding box, a CAD model, and camera intrinsics as inputs, and is licensed under the NVIDIA Open Model License for commercial use. The model is trained on synthetic data from Objaverse and Google Scanned Objects, with evaluation on LINEMOD and YCB-Video.
AIThe models were trained to grasp, place, pull, open. Then they were asked to string those moves together. No extra practice on the full task. Best score: 16.7%. Some models: ZERO. A robot can open a drawer and then get stuck on the handle.
AIOn physical Franka hardware, the success rate falls to 24%–72% of what the model achieves in simulation. With a dual-arm configuration, the range is 13%–60%. Two models that look tied in sim can be 30 points apart on hardware. Sim is the practice room. The robot is the test.
AIIntroducing FlagEval-Robo — an open, dual-track evaluation suite connecting simulation with real-world execution. We systematically post-trained and stress-tested 12 leading open-weight models under strictly aligned conditions. Here is what we discovered👇
AIRobot startups are racing to collect training data, from companies paying cleaners to wear cameras to firms recording VR-controlled humanoid robots. The article says the largest openly available robot task dataset, ABC-130K, contains only 3,500 hours of demonstrations. Skild CEO Deepak Pathak argues companies must gather high-quality data before robots can do enough useful work to generate it through deployment.
AINVIDIA released EgoHand-1.0, a 883.5M-parameter DINOv3-based transformer that predicts SOMA hand pose, MHR shape coefficients, and camera translation from a single 256×256 hand crop. The model is evaluated on the HOT3D egocentric benchmark and is intended for research and demonstration rather than production use. Its outputs can supply hand trajectories for training robotic manipulation policies, and it runs on NVIDIA Ampere GPUs under Linux with PyTorch.
AIWorld Labs has introduced Atlas, a multimodal world model it describes as trained from scratch. The post says Atlas generates frames with pixel-perfect camera control, reconstructs large scenes from as few as one input image, and outputs 3D spaces from one or more images. The author cites use cases including VFX and robotics.
AIAnthropic is building the Model Hardware Standard (MHS), a common way for AI models to connect to lab and manufacturing equipment and operate it with safety limits built into each device. MHS started as a collaboration between Anthropic and HHMI Janelia Research Campus and is launching as a research preview with partners across science, robotics, and manufacturing.
Why it matters: The source describes a standard for connecting AI models to lab and manufacturing hardware, which matters for anyone building automated experimentation workflows.
AIWould you rather live in a world filled with the same faceless white humanoids or a world where robots come in all shapes, colours and capabilities?
AIThanks @JagdeepBhatia8 for the suggestion. Read more about GEN-1.5 in our blog post in the comments below.
AIThe faster anyone can teach a robot to do something new, the easier it becomes to scale physical work. Read more about GEN-1.5 in our blog post in the comments below.