Skip to contentSkip to stories

Updated

#Open-source ecosystem

Oct 9

TodayOct 9Fri2 items
  1. QbitAIAI score64

    Tsinghua-linked VPP2 world action model tops RoboDojo simulation leaderboard

    AIStar Motion Era's VPP2, a world action model, ranked first on the RoboDojo simulation leaderboard with a 32.26% average success rate and 39.26 average score. The article attributes gains to staged training that separates video prediction from action learning, and reports a 58.5% zero-shot success rate on a real ALOHA dual-arm robot versus 40% for π0.5. The code is open source on GitHub.

  2. ArenaAI score38

    Mistral Large 4 ranks in Agent Arena top 15 at -6.6% net score

    AIMistral Large 4, a preview model from Mistral AI, ranks #43 overall in Agent Arena with a -6.6% net improvement score across more than 5,000 real-world agentic sessions. That is 11 rankings above its predecessor, Mistral Medium 3.5 (-12.60%), and places it in the top 15 labs, the only European lab there. Open weights are expected at the end of October, and at its current score the model would rank #13 among open models.

Oct 8

Oct 8Thu
  1. meng shaoAI score43

    Unsloth integrates Microsoft's mxc sandbox for Windows AI agent isolation

    AIUnsloth has integrated Microsoft's open-source mxc sandboxing system into Windows as an OS-level sandbox for isolating AI agent code execution. Its High mode provides real operating-system isolation that confines tool calls to specified directories, while its Low mode adds language-level checks that block dangerous commands and shell escapes. Both modes also strip secret environment variables and enforce resource limits such as 8GB memory and 600-second CPU time.

  2. meng shaoAI score41

    Matt Pocock shares a framework for matching AI coding workflows to change size

    AIMatt Pocock recommends matching AI coding agent workflow weight to change size: one-shot small diffs, start medium-to-large changes with /grill-with-docs for requirements clarification, and escalate to /wayfinder for mapping and tickets only when planning becomes complex. He warns against starting with /wayfinder, since a simpler-than-expected solution can leave the generated map and tickets unnecessary.

  3. meng shaoAI score77

    Theo open-sources tsc-rs, a Rust port of the TypeScript 7 compiler

    AITheo, creator of the T3 Stack, open-sourced tsc-rs, a line-by-line Rust port of Microsoft's Go-native TypeScript 7 compiler, type checker, and language server under MIT, pinned to typescript-go commit 673a5f17. The author reports tsc-rs is about 1.61× faster than tsc 7 and about 2.95× faster than bun check on six real-app benchmarks on an Apple M4 Pro. The port passes all 181,711 ported Go tests, and CLI output matches the Go version on 120 open-source repos except for known edge cases such as monorepo rootDir and tsc -b incremental output.

    Why it matters: The post reports a benchmarked, test-verified Rust port of the TypeScript 7 compiler, with pinned upstream and stated edge cases useful for judging its compatibility.

  4. PandailyAI score55

    Chinese Team Publishes 3D Cell Atlas of Rice's Full Life Cycle in Cell

    AIA Chinese-led team published in Cell a three-dimensional spatiotemporal cell atlas covering rice from germinating seed to grain fill, along with a public portal and the RICE scGPT single-cell foundation model. The atlas combines single-nucleus RNA sequencing with BGI's Stereo-seq spatial transcriptomics across 10 organ and tissue types and 61 stages, defining 119 cell types and 133 subtypes.

  5. SiliconANGLE · AIAI score62

    Nous Research raises $90M at $1.5B valuation for Hermes AI agent

    AINous Research, developer of the open-source Hermes AI agent, raised $90 million in a Series B round led by Robot Ventures, with Nvidia and Samsung participating. The company is now valued at $1.5 billion and says Hermes has been downloaded more than 24 million times. The source also reports $36 million in annualized revenue as of mid-September and expects it to exceed $100 million by year's end.

  6. The Guardian · AIAI score23

    Hollywood's Tech Ties Fuel Three Films on Zuckerberg, Musk and Altman

    AIThree upcoming films, Aaron Sorkin's The Social Reckoning, Alex Gibney's Musk, and Luca Guadagnino's Artificial, critically examine tech leaders Mark Zuckerberg, Elon Musk, and Sam Altman. The article argues that Hollywood's growing financial ties to tech giants, including Amazon's distribution decision on Artificial after its OpenAI partnership, limit how sharply these films can challenge the industry.

  7. The Guardian · AIAI score40

    Kenya's Kakuma refugees power tech microwork for dwindling, uncertain pay

    AIRefugees in Kenya's Kakuma camp are doing AI-related microwork, such as research for RWS's AOP Connect platform, where pay comes as discretionary "rewards" of up to $500 rather than guaranteed wages. Interviews with more than two dozen refugees found most earn far less than the maximum, and entry-level remote tech work is shrinking. Refugees who are unable to legally work in Kenya say they accept these gigs despite opaque pay criteria.

  8. TechCrunch · AIAI score67

    Nous Research raises $90M at $1.5B valuation to push Hermes for Businesses

    AINous Research, developer of the open source Hermes Agent, raised a $90 million Series B at a $1.5 billion valuation led by Robot Ventures. The capital will fund its enterprise push with Hermes for Businesses, which lets companies deploy customized AI agents for multi-step workflows while keeping data private. The article reports about $36 million in annualized revenue by mid-September 2026, citing The Wall Street Journal.

  9. TechCrunch · AIAI score38

    Google launches Google AI Edge Foresight, a local-first Mac meeting note-taker rival to Granola

    AIGoogle released Google AI Edge Foresight, a Mac app that captures meeting notes offline using the on-device EmbeddingGemma 2 model with 740 million parameters. The app offers split-screen shorthand and AI-generated notes, transcripts, and a Gemma 4-powered assistant that can answer questions from uploaded documents. Google's FAQ says it is optimized for Apple Silicon.

  10. Jerry LiuAI score38

    LightOn OCR-3 now on OpenDocRouter, near Gemini 3.8 Flash at lower cost

    AILightOn OCR-3 is now available on OpenDocRouter at $0.28 per 1M input tokens and $1.40 per 1M output tokens, about $3.19 per 1k pages on ParseBench. On ParseBench, the author says it sits on the Pareto frontier for open-weight OCR models, with performance similar to Gemini 3.8 Flash low at roughly 45% lower price. It is described as decent at tables, workable for charts, and quite good at grounding.

  11. Midjourney UpdatesAI score25

    Midjourney Adds Shared Folders and Thinking Mode in Alpha Update

    AIMidjourney's alpha site now lets users share folders with others as Collaborators or Viewers, with sharing by link also available. A new thinking mode lets users rerun jobs made with 8.2 standard and edit models to fix missed prompt details such as objects, layout, anatomy, and text. Sharing does not change image privacy, so non-stealth images can still appear on Explore and profiles.

  12. AnthropicAI score57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    AIAn astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  13. Tessl BlogAI score42

    Agent Skills Should Be Treated as Supply Chain Components

    AITessl's talk at AI Native DevCon London argues that agent skills, which can be markdown files with instructions and bundled material, act as supply chain components that can shape agent behavior. The author says reading SKILL.md once is insufficient because risks can sit in supporting files, updates, and workspace trust settings. He identifies the danger as the combination of private context, untrusted content, and external communication, and cites research scanning roughly 4,000 public skills for issues including malware-like behavior.

  14. Elvis SaraviaAI score42

    Don't sleep on domain-specific harnesses.

    AICoding agents are great because of their harnesses, but they aren't built for creative work. Creative work needs its own harness. Voyager looks great. It's an open harness for video, graphics, and games. The agent works with your files and drives apps like Blender, DaVinci Resolve, and Unity right on your desktop. Bring Opus, Astra, or DeepSeek. Excited to try this one.

  15. The Verge · AIAI score30

    SpaceXAI Backs Omarchy Linux Distro With $1.5 Million in Grok Tokens

    AISpaceXAI is joining the Omacom Foundation, which oversees the Omarchy Linux distribution, as a Founding Corporate Patron and donating $1.5 million worth of Grok tokens to the project. According to David Heinemeier Hansson's blog post, the tokens will primarily accelerate development, review code, and patch bugs. The partnership follows earlier controversy over Hansson's anti-immigration posts, which have drawn criticism of Omarchy's corporate contributors, including 1Password and Cloudflare.

  16. Tessl BlogAI score44

    Continuous AI Brings Agentic Automation to Repository Workflows

    AITessl's blog post argues that repository automation needs Continuous AI, a third pillar alongside CI and CD for scheduled, auditable AI workflows that improve repositories over time. The article describes GitHub Agentic Workflows, which harden agentic workflow specifications into GitHub Actions that can run coding agents such as Claude Code, Copilot CLI, Gemini CLI, or Codex-style agents. It emphasizes read-only agent steps, restricted outputs, and human review of pull requests.

  17. Goodfire ResearchAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  18. SemiAnalysisAI score38

    Open-source models absorb easier tasks, testing frontier labs' business case

    AISemiAnalysis argues that many businesses, especially low-margin ones, are offloading simpler software and white-collar tasks to increasingly capable open-source models. It frames the durability of frontier labs as depending on whether new tasks enabled by smarter frontier intelligence will outgrow the work moved to cheaper models. The post asks whether an economy could absorb 100 million superintelligent PhD-level experts quickly while still earning high ROI.