Skip to content

#Deployment/Engineering

Oct 8

TodayOct 8Thu23 items
  1. PandailyAI score41

    Simplexity Robotics Trains One Robot to Tend Two CNC Lathes With 600 Trajectories

    Simplexity Robotics says it trained a single robot to load and unload two CNC lathes on its own, using 600 real-robot trajectories and reporting a 100% success rate on the precision CNC insertion task. The work, presented at IROS 2026 on September 29, combines the SimpleWAM world action model, a DRAM memory module and DPE action scoring, with force and torque feedback for insertion recovery. The company did not say how many trials the 100% figure covers, and it does not describe the yield of the whole cell.

  2. PandailyAI score44

    Donghua University Spins Transistors Into Fibers That Act as Soft Robot Circuits

    Donghua University researchers spun transistors, resistors and capacitors into a continuous fiber that functions as a circuit, using microfluidic encoded spinning, according to a Nature Electronics paper. The fibers integrated multicolor electroluminescence, analog and digital logic, and non-contact spatial sensing, and in demonstrations guided a robotic gripper and let a finger control a robotic arm and drone without touch.

  3. PandailyAI score55

    Chinese Team Publishes 3D Cell Atlas of Rice's Full Life Cycle in Cell

    A Chinese-led team published in Cell a three-dimensional spatiotemporal cell atlas covering rice from germinating seed to grain fill, along with a public portal and the RICE scGPT single-cell foundation model. The atlas combines single-nucleus RNA sequencing with BGI's Stereo-seq spatial transcriptomics across 10 organ and tissue types and 61 stages, defining 119 cell types and 133 subtypes.

  4. Elvis SaraviaAI score55

    HERMES harness lifts GPT-5.6 Sol repository migration from 6.5% to 31.0%

    A paper introduces HERMES, a harness that pairs each repository component with a resident LLM and uses dependency-aware activation and failure diagnosis. With the same model and effort setting, GPT-5.6 Sol's whole-repository migration score rose from 6.5% to 31.0% when Codex was replaced by HERMES. Across four software engineering benchmarks, HERMES beats matched baseline harnesses by 12.4 points on average, and Qwen3-8B components come within 4.5 points of an all-GPT-5.6 Sol setup while cutting Terminal-Bench 4.0 inference cost by 26.2%.

  5. GoogleAI score62

    Google's AMIE Diagnostic Chatbot Studied Prospectively in Real Clinical Setting

    Google reports that its medical research system AMIE, described as the first patient-facing conversational diagnostic tool of its kind studied prospectively in a real-world clinical setting, was evaluated in a study published in The Lancet. Patients who chatted with AMIE before in-person appointments reported stronger confidence and better organized thoughts. Physicians reviewing the pre-visit conversation information gained more time for collaborative care and shared decision-making instead of digging through data.

  6. Epoch AIAI score31

    This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts. This is a follow-up to our earlier report on latency scaling in frontier models: https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude#appendix-h-gpt-6-sol-and-gpt-61-sol

    This isn’t conclusive evidence of a new architecture, but it suggests something has changed in how GPT-6.1 Sol handles long contexts. This is a follow-up to our earlier report on latency scaling in frontier models: https://epoch.ai/publications/long-context-latency-scaling-gpt-vs-claude#appendix-h-gpt-6-sol-and-gpt-61-sol

  7. Elvis SaraviaAI score48

    Google's FlowAgent auto-repairs failing tests inside code review

    Google proposed FlowAgent, a ReAct-style agent that generates and validates fixes for pre-submit test failures and shows them in its code review tools. Two abstention filters, before and after execution, suppress weak suggestions; in a manual review of 195 real failures, 67.18% of fixes were correct. After the Google-wide launch, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554.

  8. GoodfireAI score28

    We built cybersecurity monitors for Kimi K3 and GLM 5.3 — more accurate, 50x faster, & 50x cheaper than an optimized LLM judge. External red-teaming from @farairesearch found that our monitor greatly reduces universal jailbreaks. 🧵

    We built cybersecurity monitors for Kimi K3 and GLM 5.3 — more accurate, 50x faster, & 50x cheaper than an optimized LLM judge. External red-teaming from @farairesearch found that our monitor greatly reduces universal jailbreaks. 🧵

  9. Goodfire ResearchAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    Goodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  10. ZyphraAI score34

    Our approach is lossless. The same tokens visit the same experts. The architecture, routing decisions and training objective stay unchanged. The model computes the same function. By organizing where experts and tokens live we reduce the communication needed to do the same work.

    Our approach is lossless. The same tokens visit the same experts. The architecture, routing decisions and training objective stay unchanged. The model computes the same function. By organizing where experts and tokens live we reduce the communication needed to do the same work.

  11. ZyphraAI score38

    The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster

    The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster

  12. ZyphraAI score18

    Routing is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling moves each token to the GPU holding those experts, inside a transfer that already runs after attention, so it adds no network traffic.

    Routing is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling moves each token to the GPU holding those experts, inside a transfer that already runs after attention, so it adds no network traffic.

  13. ZyphraAI score37

    So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.

    So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.

  14. ZyphraAI score32

    MoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and tokens have to be sent to their experts and back. In our runs that exchange took 13-24% of step time on one node and 45-60% across four nodes.

    MoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and tokens have to be sent to their experts and back. In our runs that exchange took 13-24% of step time on one node and 45-60% across four nodes.

  15. ZyphraAI score23

    Those routing decisions have patterns. For instance, in one layer, just 64 of 8,128 possible expert pairs accounted for 42% of tokens. If experts were picked independently, those pairs would carry 1.6%. We show how to exploit this structure to improve where experts are placed.

    Those routing decisions have patterns. For instance, in one layer, just 64 of 8,128 possible expert pairs accounted for 42% of tokens. If experts were picked independently, those pairs would carry 1.6%. We show how to exploit this structure to improve where experts are placed.

  16. ZyphraAI score22

    In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime. At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.

    In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime. At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.

  17. QbitAI (量子位)AI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    UniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  18. PandailyAI score38

    Huawei Presents Experimental XMFS Shared-Memory Filesystem at LPC 2026

    Huawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.

Oct 7

Oct 7Wed
  1. vLLMAI score46

    vLLM-Omni technical report unifies serving for omni-modality generation

    The vLLM team released a technical report on vLLM-Omni, a unified serving runtime for omni-modality generation spanning multi-stage autoregressive pipelines, iterative diffusion, and stateful sessions. Current LLM servers and diffusion stacks each cover only one of these patterns, pushing deployments to stitch disjoint runtimes together. vLLM-Omni offers a shared control plane in which an orchestrator advances requests across stages, specialized engines handle compute, and a connector carries payloads.

  2. Google ResearchAI score42

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

  3. Google ResearchAI score62

    Google Research finds AI boosts patent drafting but junior lawyers' gains vanish without it

    A Google Research field experiment with 133 patent lawyers found AI tool access raised drafting scores by 0.34 to 0.38 standard deviations over three months. When the tool was removed for a redlining task, only senior lawyers kept an advantage of 0.45 SD, while junior lawyers showed no discernible improvement. The authors argue that tools which boost current output must not stop junior professionals from building the judgment that senior experts rely on.

    AIWhy it matters: The field experiment separates AI's short-term productivity gains from skill retained after the tool is removed, which matters for training junior professionals.

  4. Wired · AIAI score60

    Researchers Test GPT-6 Astra Driving a Corolla to In-N-Out

    Three Axiom engineers had OpenAI's GPT-6 Astra drive a 2024 Toyota Corolla to an In-N-Out drive-thru through a server linked to cameras and power steering, with a safety driver ready to brake. They also built a parking-lot benchmark, DrivingBench, where Astra completed the course slowly, Claude Fable 5.1 finished 45 percent, and Grok finished 11 percent.

  5. WaymoAI score27

    The best time to prepare for an emergency is before it happens. New Waymo research introduces a first-of-its-kind framework for AV incident-management exercises—from tabletop scenarios to full-scale drills. Adapted from emergency management best practices, the framework helps AV developers, operational partners, and first responders test plans and strengthen coordination together. Read more: https://waymo.com/blog/2026/10/incident-management-exercises/

    The best time to prepare for an emergency is before it happens. New Waymo research introduces a first-of-its-kind framework for AV incident-management exercises—from tabletop scenarios to full-scale drills. Adapted from emergency management best practices, the framework helps AV developers, operational partners, and first responders test plans and strengthen coordination together. Read more: https://waymo.com/blog/2026/10/incident-management-exercises/

  6. Elvis SaraviaAI score44

    NVIDIA's VERA co-evolves agent harness and model via verifiable environments

    NVIDIA's VERA turns benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring and updates both model weights and the agent harness together. A harness edit is kept only if it adds at least 5 points on the development set, and a checkpoint is rejected if its score drops more than 20%. At 27B, the co-evolved agent scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and the environment corpus is open-sourced.

Oct 6

Oct 6Tue
  1. Waymo BlogAI score31

    Waymo Publishes Framework for Autonomous Vehicle Incident Management Exercises

    Waymo researchers and incident readiness experts published a paper introducing a framework to help AV developers plan, test and strengthen incident-management capabilities. The framework adapts FEMA's Homeland Security Exercise and Evaluation Program for automated vehicle operations and outlines four exercise types: formative, educational, summative and confirmatory.

  2. PyTorch BlogAI score46

    PyTorch Introduces FBTriton Kernels to Speed Table Batched Embedding Operations

    PyTorch's blog describes a Triton-based implementation of Table Batched Embedding (TBE) forward and backward kernels for recommendation-system embedding lookups, which the post says outperforms legacy CUDA kernels on these workloads. On B200, an updated CUDA bounds-check step reaches up to 1.24x speedup on that component, and an optional forward-side preprocessing path cuts combined latency from 79.537 ms to 66.183 ms (−16.8%) on a large configuration.

  3. OpenAIAI score62

    OpenAI releases new mathematical results from an internal frontier model

    OpenAI is releasing a broad range of new mathematical results produced by an internal frontier model. The company says it consulted the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study and drew on its advice and public recommendations for how the results are released. The results are available at https://github.com/openai/math.

  4. GoogleAI score30

    Google Earth AI helps forecast disease outbreak spread faster

    Google Earth AI, according to new research, can help communities respond to public health crises more quickly and proactively. The post says it combines behavioral trends, geospatial AI models, and other insights beyond simple statistics to help public health teams understand complex issues and bridge reporting gaps. The aim is to shift emergency response from reactive management toward proactive prevention.

  5. MIT News · AIAI score23

    MIT Lincoln Lab's LAICS Survey Tracks AI Accelerator Performance and Power Trends

    The Lincoln Laboratory Supercomputing Center's Lincoln AI Computing Survey (LAICS) has been comparing commercial AI accelerators by peak performance and peak power since 2018. The latest paper covers more than 120 accelerators, up from 57 in the first, with data drawn from public sources. The team says five to 10 new AI accelerator startups emerge each year, and six have announced their first accelerators in recent months.

  6. Google · Innovation & AIAI score42

    Google Study Tests AI-Guided Blind Sweep Ultrasounds for Pregnant Women in Kenya and Chicago

    Google researchers, working with Northwestern Medicine and Jacaranda Health, trained healthcare workers to perform "blind sweep" ultrasounds analyzed by machine learning models. The models estimated gestational age and fetal presentation as accurately as a trained sonographer in a study of 1,000 mothers each in Nairobi and Chicago. The AI processes results on the device, so it needs no electricity supply or Wi-Fi.

  7. METRAI score36

    Here’s a simple example: a METR researcher found a bug in the transcript viewer of Inspect, a popular evaluation framework, that would have enabled an agent to show the user reviewing its transcript a fake (or edited) version.

    Here’s a simple example: a METR researcher found a bug in the transcript viewer of Inspect, a popular evaluation framework, that would have enabled an agent to show the user reviewing its transcript a fake (or edited) version.

  8. METRAI score40

    In order to notice when AI agents misbehave, AI companies often log the actions and reasoning steps their agents take. However, misaligned AI agents may be able to hack the software that humans use to review and understand these logs, hiding misbehavior.

    In order to notice when AI agents misbehave, AI companies often log the actions and reasoning steps their agents take. However, misaligned AI agents may be able to hack the software that humans use to review and understand these logs, hiding misbehavior.

  9. Google ResearchAI score51

    Google's PDFM location embeddings improve five global public health tasks

    Google Research reports that Population Dynamics Foundation Model (PDFM) embeddings, built from search trends, mobility, built environment, and weather signals, were tested by partners across five public health tasks. The embeddings improved results in cross-border MMR vaccination coverage, dengue forecasting, postpartum depression screening, and cholera outbreak prediction, and matched census inputs for cardiovascular mortality nowcasting.

  10. Sophia YangAI score26

    Reinforcement learning at scale: - Autoscaling actor fleet → tens of thousands of rollouts in parallel, async training - Built for long trajectories: millions of tokens per rollout, multiple compactions, low staleness - New methods at both stages cut off-policy drift → stable long-horizon RL - 3k GPUs → ~33B tokens/day, ~16B trainable after filtering + masking - Rewards climb across representative envs as the policy learns harder tasks ↓

    Reinforcement learning at scale: - Autoscaling actor fleet → tens of thousands of rollouts in parallel, async training - Built for long trajectories: millions of tokens per rollout, multiple compactions, low staleness - New methods at both stages cut off-policy drift → stable long-horizon RL - 3k GPUs → ~33B tokens/day, ~16B trainable after filtering + masking - Rewards climb across representative envs as the policy learns harder tasks ↓

Oct 5

Oct 5Mon
  1. Chips and CheeseAI score45

    NVIDIA's Olympus Core Pushes Server Single-Threaded Performance Boundaries

    NVIDIA's Olympus is a 10-wide out-of-order server core running at 3.3 GHz that prioritizes per-clock performance over high clock speeds. It uses a simultaneous multi-threading (SMT) implementation, unlike Arm's Cortex X925, and has out-of-order structures larger than X925's. In SPEC CPU2026, its branch prediction accuracy is slightly behind AMD's Zen 5 and slightly ahead of Intel's Lion Cove.