Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. Midjourney UpdatesOfficialAI score25

    Midjourney Adds Shared Folders and Thinking Mode in Alpha Update

    AIMidjourney's alpha site now lets users share folders with others as Collaborators or Viewers, with sharing by link also available. A new thinking mode lets users rerun jobs made with 8.2 standard and edit models to fix missed prompt details such as objects, layout, anatomy, and text. Sharing does not change image privacy, so non-stealth images can still appear on Explore and profiles.

  2. Hacker News · Show HN, AI (20+ points)BlogAI score62

    Rembrandt releases a free, open-source local photo editor with on-device AI

    AIRembrandt is a free, open-source photo editor for macOS, Windows and Linux that runs locally and needs no account or subscription. It saves edits as standard XMP sidecar files next to each photo and can also run as a self-hosted server accessed through a browser. Only optional cloud sync is paid.

  3. AnthropicOfficialAI score57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    AIAn astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  4. Claude Code · GitHub ReleasesOfficialAI score56

    Claude Code v2.1.295 adds hook failure blocking and gateway controls

    AIClaude Code v2.1.295 adds onFailure: "block" for command and HTTP hooks, so a hook that cannot start, times out, or exits unexpectedly blocks the action. The release also adds an optional models list for Claude apps gateway upstreams, plus upstream_request_id in the inference audit event, and fixes a range of MCP, plugin, and terminal issues.

  5. Sherwin WuXAI score60

    Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%

    AISherwin Wu, an OpenAI employee, says the updated Harvey LAB-AA v1.1 benchmark, announced by Artificial Analysis with Harvey, is more useful than the original LAB results. The new Hallucination-Gated All-Pass Rate credits a task only when every rubric criterion passes and no material hallucination appears. Grok 4.7 (xhigh) leads at 9.4%, while GPT-6 Astra (max) at 8.6% has very few material hallucinations.

    Why it matters: The update adds a hallucination gate to a legal benchmark, showing that models with high all-pass rates can rank much lower once material errors count.

  6. Codex · GitHub ReleasesOfficialAI score36

    Codex 0.162.0 adds managed worktree tools and clickable URLs in the TUI

    AIOpenAI's Codex 0.162.0 release adds tools for creating and listing managed Git worktrees from trusted local projects when the worktrees feature is enabled. The update also lets users pin tasks in the agent Command Center, copy transcript blocks with /copy, and make URLs clickable in approval headers, questions, and warnings, along with several Linux and Windows sandbox fixes.

  7. Tessl BlogOfficialAI score42

    Agent Skills Should Be Treated as Supply Chain Components

    AITessl's talk at AI Native DevCon London argues that agent skills, which can be markdown files with instructions and bundled material, act as supply chain components that can shape agent behavior. The author says reading SKILL.md once is insufficient because risks can sit in supporting files, updates, and workspace trust settings. He identifies the danger as the combination of private context, untrusted content, and external communication, and cites research scanning roughly 4,000 public skills for issues including malware-like behavior.

  8. elvisXAI score42

    Voyager: an open harness for creative AI work across video and games

    AIElvis Saravia argues that creative work needs domain-specific agent harnesses rather than coding-oriented ones, and he highlights Voyager as an open harness for video, graphics, and games. According to the quoted post, Voyager lets agents work with local files and drive apps such as Blender, DaVinci Resolve, and Unity, and it is designed to work with models like Opus, Astra, and DeepSeek.

    Video from @omarsar0's post
  9. Tessl BlogOfficialAI score38

    Mozilla.ai's cq Aims to Give Agents a Shared, Reviewable Knowledge Commons

    AIMozilla.ai's cq project proposes a shared knowledge layer where AI agents capture lessons from non-obvious fixes as structured knowledge units that other agents can later query. The default setup is local-first, using a local SQLite database so nothing leaves the machine, with an option to connect to a remote team server that adds review.

  10. Artificial AnalysisOfficialAI score28

    Artificial Analysis Pareto frontier: GPT-6 Luna cheapest per task at $0.22

    AIAmong models with a Hallucination-Gated All-Pass Rate above 0%, GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max), and Grok 4.7 (xhigh) set the Pareto frontier for score versus cost per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%, while Grok 4.7 (xhigh) leads at about $9.50 per task and Muse Spark 1.3 (max) costs about $4.20. The three Claude models cost about $18 to $22 per task.

    Image from @ArtificialAnlys's post
  11. DatabricksOfficialAI score32

    Databricks' Vibe Data Modeling builds business-specific data models with an agent

    AIDatabricks introduced Vibe Data Modeling, an open-source agent that helps teams build, validate, and evolve business-specific data models. It applies roughly 250 modeling rules while keeping data modelers and business stakeholders involved. Teams can start from 40 industry models as a baseline and iterate toward models that reflect how their business operates.

    Video from @databricks's post
  12. laurenXAI score29

    Omarchy seeks feedback on Grok Bot plugins and integrations

    AILauren Tan invites users of Grok Bot on Omarchy and developers building plugins for it to share feedback and feature requests. The post points to the Omarchy plugin catalog and asks what integrations could be supported. Background from DHH says SpaceXAI joined the Omacom Foundation as a Founding Corporate Patron, contributing $1,500,000 in Grok tokens for Omarchy's maintenance and development.

  13. PyTorch BlogOfficialAI score62

    NVIDIA Dynamo adds session-level IDs to route and cache agentic inference

    AINVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.

    Why it matters: The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.

  14. KalaXAI score34

    Mistral Large 4 and Reflection Beam promise open weights this month

    AIMistral Large 4 and Reflection Beam are previewed now, with Mistral saying weights drop at the end of October and Reflection promising Apache 2.0 weights this month. The post argues that these announced future weights should be treated as a conditional migration dependency, not a current self-hosting option. API previews can be trialed immediately, but they do not prove an unreleased checkpoint will behave the same when downloaded.

  15. Hacker News · Show HN, AI (20+ points)BlogAI score43

    Pocketty is an iPhone SSH terminal that alerts you when an agent is blocked

    AIPocketty is a $99 iPhone and iPad SSH terminal, with a 14-day free trial, that notifies you when an herdr-managed agent is blocked or done. The alert is sealed on your computer for your phone only, and tapping it opens that exact Pane over SSH so you can answer in a real terminal. The source says the relay forwards only sealed bytes and that terminal traffic goes directly between the app and your computers.

  16. The Verge · AINewsAI score30

    SpaceXAI backs Omarchy Linux distro with $1.5 million in Grok tokens

    AISpaceXAI joins the Omacom Foundation, which oversees the Omarchy Linux distro, as a Founding Corporate Patron and donates $1.5 million in Grok tokens. According to David Heinemeier Hansson's blog post, the tokens will primarily accelerate development, review code, and patch bugs. The article notes Hansson's recent anti-immigration blog posts, and that 1Password and Cloudflare have also faced criticism for contributing to Omarchy.

  17. Alexander DoriaXAI score46

    LightOnOCR-3 claims state-of-the-art OCR performance under 1B parameters

    AILightOn has released LightOnOCR-3, a family of OCR models in 0.8B and 4B versions that it says lead benchmarks including OlmOCR-Bench and ParseBench, with the 0.8B model positioned as the sub-1B option. The models recognize text, handwriting, images, charts and document structure in one pass, process documents up to twice as fast as LightOnOCR-2, and are released under the Apache 2.0 license.

    Image from @Dorialexander's post
  18. Dhravya ShahXAI score42

    MemoryRepo: open-source implementation of Cognition's dreaming agent memory

    AISupermemory introduces MemoryRepo.dev, an open-source implementation of Cognition's dreaming memory system built on Cloudflare Artifacts, Durable Objects, Alchemy, and Effect. The project follows Cognition's Devin memory design, which builds a memory graph across sessions and prunes stale records overnight. Supermemory says it will incorporate learnings from this research into its own product.

    Video from @DhravyaShah's post
  19. Zhihao JiaXAI score62

    Lithos AI open-sources lithos-metal for fast local inference on Apple M5 Max

    AILithos AI says it is open-sourcing lithos-metal, which uses megakernels and DSpark speculative decoding. The post claims Qwen3.8-27B reaches a peak of over 200 tokens per second per user on a single Apple M5 Max. It says users can try the tool with any coding agent in one command, and links to the code on GitHub and a technical blog.

    Video from @JiaZhihao's post
  20. ClineOfficialAI score46

    Cline makes Step 5 Preview free, citing strong DeepSWE coding scores

    AICline says Step 5 Preview is now free in its coding tool and scores ahead of Kimi K3 and GLM-5.3 on DeepSWE. The company describes it as one of the strongest open-weights coding models available. StepFun's background announcement describes Step 5 Preview as a 600B total / 27B active MoE model with 1M context and vision, and says open weights arrive on Oct 15.

    Image from @cline's post
  21. Hacker News · Show HN, AI (20+ points)BlogAI score43

    Show HN: AI SRE Arena, an open benchmark for AI SRE agents on Kubernetes

    AIAI SRE Arena is an open, vendor-neutral benchmark that injects faults into a disposable Kubernetes fixture and scores AI SRE investigations with a configurable judge. In its first published comparison across 21 incident scenarios, Edge Delta's native AI investigations detected 18 of 21 incidents, and Grafana's detected 12. Claude, run through each vendor's observability CLI, reached 90.5% root cause analysis accuracy with Grafana's CLI and 85.7% with Edge Delta's.

  22. Tessl BlogOfficialAI score44

    Continuous AI Brings Agentic Automation to Repository Workflows

    AITessl's blog post argues that repository automation needs Continuous AI, a third pillar alongside CI and CD for scheduled, auditable AI workflows that improve repositories over time. The article describes GitHub Agentic Workflows, which harden agentic workflow specifications into GitHub Actions that can run coding agents such as Claude Code, Copilot CLI, Gemini CLI, or Codex-style agents. It emphasizes read-only agent steps, restricted outputs, and human review of pull requests.

  23. elvisXAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  24. Hacker News · Show HN, AI (20+ points)BlogAI score23

    Show HN: Jevman lets AI models play Pac-Man against the arcade ghosts

    AIJevman is an open-source Pac-Man benchmark where AI models play 100 games each against the classic scripted ghosts. Each model gets a maze state at every junction and returns a direction probability, with answers over 2 seconds replaced by a backup rule. Community models can join the leaderboard by submitting games that CI replays to verify their scores.

  25. Daniel HanXAI score38

    Unsloth adds OS-level sandboxing for Linux, Mac, and Windows

    AIUnsloth now supports OS-level sandboxing on Linux via bwrap, on Mac via seatbelt, and on Windows via Microsoft's MXC. Per-tool-call latency is under 100ms across all three, and its software-style sandboxing with regex AST checks adds about 3ms. The Windows integration was built in collaboration with Microsoft.

  26. SunoOfficialAI score22

    Suno launches Albums for bundling songs into full releases

    AISuno announced that Albums are now live, letting users combine songs into a full release, set artwork, arrange the tracklist, and publish when ready. Existing playlists can be converted into Albums without rebuilding them from scratch.

    Video from @suno's post
  27. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  28. Unsloth AIOfficialAI score44

    Unsloth adds Windows OS-level sandboxing via Microsoft's mxc

    AIUnsloth now supports OS-level sandboxing on Windows by integrating Microsoft's open-source mxc repository for sandboxed code execution. The integration adds under 100 ms of overhead, according to the post. A setup guide is available in Unsloth's documentation.

    Image from @UnslothAI's post
  29. Goodfire ResearchOfficialAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.