Skip to contentSkip to stories

Updated

#Open-source ecosystem

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. Orange AIAI score34

    Next Token episode 5 covers Personal Agents, open-source software, and hardware projects

    AIThis Next Token episode discusses Personal Agents, including Dots in Codex, memory and cloud computer permissions, and whether agents should act as assistants or digital twins. The hosts also cover Instinct's booking and business-travel model, hands-on projects built with Opus 5.5, and whether software, games, and hardware could become open source as AI makes rewriting easier.

  2. MarkTechPostAI score58

    Unsloth Studio re-checks changed model repos and blocks flagged weights before loading

    AIUnsloth Studio binds remote-code approval to a fingerprint of the scanned code, so changed code requires fresh consent before it runs. It also blocks weight files that Hugging Face has flagged for malware in the path the selected loader would deserialize. The article describes these checks as one layer among several, alongside package-content scans and OS sandboxes, and notes that the scanner is not a sandbox and cannot catch every evasion.

  3. InferactAI score38

    Inferact and partners cut vLLM TTFT nearly 70% at ~100K throughput

    AIInferact, working with DeepSeek, NVIDIA, and SemiAnalysis alongside the vLLM community, says joint work across models, custom kernels, and engine serving cuts time to first token (TTFT) by nearly 70% at ~100K throughput. vLLM is the open-source inference engine, and Inferact optimizes it for enterprise production deployments.

  4. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  5. Hugging Face BlogAI score66

    How one developer built six custom models with ML-Intern for about USD 103

    AIA Hugging Face blog author used the ML-Intern agent in HuggingChat to build six small models by writing detailed prompts that specify datasets, base models, baselines, smoke tests, and spending limits. The projects include a citrus disease vision-language model, a Huggy character LoRA, a camera-angle LoRA, a doodle-to-object LoRA, a 0.8B prompt rewriter, and a 4-step distilled Agate model, with total compute cost of about USD 103. Each project's prompts and public models are linked from the post.

    Why it matters: The author shows how prompt structure, baselines, smoke tests, and budget caps shape an agent-driven training workflow, with per-project costs given.

  6. Ars Technica · AIAI score46

    Artcraft releases open source clones of Adobe Photoshop, Premiere and other apps built with Claude

    AIDeveloper Brandon Thomas's Artcraft has launched seven open source apps in Rust that aim to replicate the interfaces and tools of Adobe Photoshop, Illustrator, Premiere, Lightroom, After Effects, InDesign, and Acrobat Pro. Thomas said he used Anthropic's Claude Opus 5.5 to build the clean-room replacements, with WebAssembly versions available for browser use. The apps remain in a "super early alpha" state, and commenters have pointed out many current shortcomings.

  7. Leandro von WerraAI score36

    Snorkel expands Open Benchmarks Grants to $30M for AI evaluation

    AISnorkel AI is expanding its Open Benchmarks Grants tenfold to a $30M commitment to fund more diverse, robust, and continuously updated open AI benchmarks. The program adds an Open Benchmarks Red Team to test and strengthen those benchmarks, plus a Snorkel Research Fellowship for independent researchers developing new evaluation methods. The source says OBG-funded benchmarks have appeared on model cards from every major frontier lab.

  8. OpenRouterAI score46

    Perplexity Decider v1.1 arrives on OpenRouter with free output

    AIPerplexity's open-weights multimodal decision model, Decider v1.1, is now available on OpenRouter. It accepts text, JSON, or images and returns typed answers with probabilities, priced at $0.02 per million input tokens with free output. Perplexity says it scores highest on Hugging Face's new Decision Index 0.3 benchmark and costs half as much as v1.

  9. MarkTechPostAI score60

    Liquid AI releases open-weight d1-3B and d1-omni-600M decision models

    AILiquid AI released two open-weight multimodal decision models, d1-3B and d1-omni-600M, which return probability answers in one forward pass with zero output tokens. d1-3B scores 48.57 on Decision Index v0.2.1 and answers one question in 8 ms on an RTX 4090, while the models are licensed free for commercial use below $10 million in annual revenue.

  10. Georgi GerganovAI score44

    llama.cpp adds ggml RPC for distributing inference across heterogeneous devices

    AIllama.cpp can distribute inference across heterogeneous devices through the ggml RPC backend, according to Georgi Gerganov. He says it is currently an advanced setting, but he expects it to become more accessible to regular users over time. A related post reports MiMo 2.6 Flash running across an RTX 6000 GPU and an M5 laptop over 10 GbE at about 40 tokens/sec.

  11. LlamaIndex 🦙AI score47

    LlamaIndex launches OpenDocRouter, one API for many document parsing models

    AILlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.

    Video from @llama_index's post
  12. Demis HassabisAI score34

    Isomorphic Labs joins Virtual Biology Initiative as founding member

    AIIsomorphic Labs has joined the Virtual Biology Initiative (VBI) as a founding member, helping build an open resource for global researchers to accelerate understanding and treatment of disease. Demis Hassabis, owner of the account and Google/Gemini-affiliated, called applying AI to medicine and human health the most important thing AI can be used for and welcomed the partnership with CZI and VBI.

  13. Lucas Beyer (bl16)AI score36

    Reality Check: a public leaderboard for robot manipulation VLA models

    AILucas Beyer praises Reality Check, a new leaderboard for benchmarking VLA and related robot manipulation models. Half of its tasks are fully open, while the other half are held out to detect benchmaxxing by future model versions. The companion post from Nicolas Keller describes the launch as the first public robot manipulation benchmark, built on 14,400 real-world rollouts across four models.

  14. GitHub Copilot ChangelogAI score42

    GitHub Copilot CLI adds discovery of local Ollama models via /model

    AIGitHub Copilot CLI version 1.0.94-0 lets users run /model to discover supported models from a running local Ollama instance alongside configured and GitHub Copilot cloud models. Discovered models are not added automatically; users choose one, review its provider and endpoint, then confirm Add and use for this session or Add without switching, and models must support tool calling and streaming. Choosing a local model does not enable offline mode or disable GitHub telemetry, and COPILOT_OFFLINE=true remains a separate explicit setting.

  15. elvisAI score44

    NVIDIA's VERA co-evolves agent harness and model via verifiable environments

    AINVIDIA's VERA turns benchmark trajectories into over 9,000 restartable sandboxes with rubric scoring and updates both model weights and the agent harness together. A harness edit is kept only if it adds at least 5 points on the development set, and a checkpoint is rejected if its score drops more than 20%. At 27B, the co-evolved agent scores 71.6 on AutoCoWorkBench, above Claude Opus 4.8, and the environment corpus is open-sourced.

    Image from @omarsar0's post