Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. Hugging Face BlogOfficialAI score66

    How one developer built six custom models with ML-Intern for about USD 103

    AIA Hugging Face blog author used the ML-Intern agent in HuggingChat to build six small models by writing detailed prompts that specify datasets, base models, baselines, smoke tests, and spending limits. The projects include a citrus disease vision-language model, a Huggy character LoRA, a camera-angle LoRA, a doodle-to-object LoRA, a 0.8B prompt rewriter, and a 4-step distilled Agate model, with total compute cost of about USD 103. Each project's prompts and public models are linked from the post.

    Why it matters: The author shows how prompt structure, baselines, smoke tests, and budget caps shape an agent-driven training workflow, with per-project costs given.

  2. Teknium 🪽XAI score28

    Teknium posts a "hello" greeting on X

    AITeknium, a Nous Research affiliate, posted only the word "hello" on X. The quoted post from Nous Research announces a Series B raise to advance Hermes Agent and build a mobile app, with investors including NVIDIA and Samsung Next.

  3. Ars Technica · AINewsAI score46

    Artcraft releases open source clones of Adobe Photoshop, Premiere and other apps built with Claude

    AIDeveloper Brandon Thomas's Artcraft has launched seven open source apps in Rust that aim to replicate the interfaces and tools of Adobe Photoshop, Illustrator, Premiere, Lightroom, After Effects, InDesign, and Acrobat Pro. Thomas said he used Anthropic's Claude Opus 5.5 to build the clean-room replacements, with WebAssembly versions available for browser use. The apps remain in a "super early alpha" state, and commenters have pointed out many current shortcomings.

  4. Perplexity DevelopersOfficialAI score40

    Perplexity's pplx-decider-v1.1-27b now available on OpenRouter

    AIPerplexity's open-weights multimodal decision model, pplx-decider-v1.1-27b, is now available on OpenRouter. It accepts text, JSON, or images and returns typed answers with probabilities, priced at $0.02 per million input tokens with free output.

  5. Leandro von WerraXAI score36

    Snorkel expands Open Benchmarks Grants to $30M for AI evaluation

    AISnorkel AI is expanding its Open Benchmarks Grants tenfold to a $30M commitment to fund more diverse, robust, and continuously updated open AI benchmarks. The program adds an Open Benchmarks Red Team to test and strengthen those benchmarks, plus a Snorkel Research Fellowship for independent researchers developing new evaluation methods. The source says OBG-funded benchmarks have appeared on model cards from every major frontier lab.

  6. OpenRouterOfficialAI score46

    Perplexity Decider v1.1 arrives on OpenRouter with free output

    AIPerplexity's open-weights multimodal decision model, Decider v1.1, is now available on OpenRouter. It accepts text, JSON, or images and returns typed answers with probabilities, priced at $0.02 per million input tokens with free output. Perplexity says it scores highest on Hugging Face's new Decision Index 0.3 benchmark and costs half as much as v1.

  7. vLLMOfficialAI score22

    vLLM and NVIDIA cut TTFT nearly 70% at ~100K throughput

    AITogether with DeepSeek's model and kernels and NVIDIA's collaboration, vLLM reports that time to first token (TTFT) drops nearly 70% at roughly 100K throughput. The post credits Inferact and the vLLM community for the work and thanks SemiAnalysis for AgentX.

  8. KushXAI score23

    Fluffles open-sourced as a lesson on stateful server-based agents

    AIDeveloper team open-sources fluffles-os on GitHub, presenting it as a hard lesson rather than the product at puffle.ai. The post says stateful server-hosted agents like Hermes proved unworkable, and the team's earlier fluffles agent, built as a near-unrestricted "god agent" on a Mac mini, was painful to harness because failure modes were unbounded. The team says it later ported to Eve, which let them focus on agent behavior instead of integration scaffolding, and plans a launch this week.

    Image from @kushbhuwalka's post
  9. GitHubOfficialAI score57

    GitHub Copilot local sandboxing becomes generally available

    AILocal sandboxing for GitHub Copilot is now generally available. It lets Copilot run commands in an isolated environment with controlled access to files, networks, system capabilities, and credentials. Enterprise teams can also centrally manage policies, and the feature is available in GitHub Copilot CLI, the GitHub Copilot app, and @code.

  10. TinkerOfficialAI score31

    IdeaLens detects whether ideas originated from humans or AI

    AIIdeaLens is a detector that identifies whether the ideas in a text came from a human or an AI, rather than judging the prose alone. On mixed-provenance benchmarks, it reached 81.3% average idea-detection accuracy, versus 25.4% for ProseLens and 25.9% for Pangram 4. The model, code, and data are open-sourced, and it was trained on Tinker.

  11. 👩‍💻 Paige BaileyOfficialAI score36

    Google's EmbeddingGemma 2 model released on Hugging Face for science

    AIGoogle released EmbeddingGemma 2 on Hugging Face, and Paige Bailey called it a step toward open models for open science. The post links the model and cites earlier EmbeddingGemma-based projects, including medical, geographic, oncology, and PubMed embedding models.

  12. MarkTechPostNewsAI score60

    Liquid AI releases open-weight d1-3B and d1-omni-600M decision models

    AILiquid AI released two open-weight multimodal decision models, d1-3B and d1-omni-600M, which return probability answers in one forward pass with zero output tokens. d1-3B scores 48.57 on Decision Index v0.2.1 and answers one question in 8 ms on an RTX 4090, while the models are licensed free for commercial use below $10 million in annual revenue.

  13. OpenRouterOfficialAI score38

    Cloudflare's Clef decision models now available on OpenRouter

    AICloudflare's open-source Clef (27B) and Clef Flash (9B) decision models are available on OpenRouter. They accept text, JSON, or images and return typed answers with probabilities rather than generated text. Pricing is $0.24 per M input tokens for Clef and $0.09 per M for Clef Flash, with output free.

  14. Claude Code · GitHub ReleasesOfficialAI score36

    Claude Code v2.1.293 adds Claude Haiku 5.5 and fixes dozens of bugs

    AIClaude Code v2.1.293 adds Claude Haiku 5.5 (claude-haiku-5-5), now the default Haiku model on the Anthropic API, with 1M context and pricing of $0.10/$0.50 per Mtok ($0.50/$2.50 for prompts over 100K). The release also adds agentType to the subagentStatusLine payload and isDeferred to $.tool.register, and fixes numerous issues including a memory leak in HTTP MCP connections.

  15. a16z NewsBlogAI score46

    a16z backs Preference Model, which builds RL environments for training AI models

    AIPreference Model is open-sourcing Karotte, the framework it uses to build reinforcement learning environments that resist reward hacking, including defenses like killing stray processes before grading and rejecting grader-crashing files. The framework has been hardened through more than a million evaluation runs and controlled red-teaming. The company focuses on machine learning engineering tasks for leading labs, and a16z says it is partnering with Preference Model and its founders, Jennifer Zhou and Ning Cao.

  16. Georgi GerganovXAI score44

    llama.cpp adds ggml RPC for distributing inference across heterogeneous devices

    AIllama.cpp can distribute inference across heterogeneous devices through the ggml RPC backend, according to Georgi Gerganov. He says it is currently an advanced setting, but he expects it to become more accessible to regular users over time. A related post reports MiMo 2.6 Flash running across an RTX 6000 GPU and an M5 laptop over 10 GbE at about 40 tokens/sec.

  17. Hacker News · Show HN, AI (20+ points)BlogAI score42

    Terse: a Claude Code plugin that halves reply length by cutting filler

    AITerse, a Claude Code plugin, enforces a shorter reply style by cutting filler, slogans, metaphors and first-person narration. The author says it reduces replies by 46% on Fable 5.1 and 54% on Opus 5.5, measured on 20 prompts against a real codebase. The plugin installs from the Claude directory or GitHub, and its rules apply to replies, documents, commits and subagents.

  18. Semafor · TechnologyNewsAI score56

    Reflection AI and Mistral launch open models to challenge China's lead

    AIReflection AI and Mistral each unveiled new open-source models this week, aiming to beat other Western open models, though they trail top Chinese and closed systems on prominent benchmarks. Reflection CEO Misha Laskin says the target is regulated industries and governments that cannot or will not use Chinese models. The outcome depends on whether businesses and agencies accept less advanced models for some tasks in exchange for lower cost and more control.

  19. Liquid AIOfficialAI score38

    Liquid AI's Open d1 models run on NVIDIA hardware with llama.cpp support

    AILiquid AI's Open d1 models run across NVIDIA DGX, RTX, and Jetson hardware, with day-one llama.cpp support for deployment anywhere. Measured one request at a time, the d1-3B model's single-question latency is 8 ms on an NVIDIA RTX 4090, 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin 64 GB, and 50 ms on Jetson Orin Nano.

  20. Liquid AIOfficialAI score23

    Liquid AI's d1-3B tops sub-10B models on Decision Index v0.2.1

    AILiquid AI's d1-3B ranks first among models under 10B parameters on the Decision Index v0.2.1, a benchmark for structured decision-making. Built from LFM2.5-VL-3B, it makes decisions from text and images in a single pass. It is suited to reranking, agent guardrails, and visual inspection.

    Image from @liquidai's post
  21. Liquid AIOfficialAI score52

    Liquid AI releases open-weight d1-3B and d1-omni-600M multimodal models

    AILiquid AI released Open d1, two open-weight multimodal models in its d1 decision model family. The d1-3B model supports text and vision, while d1-omni-600M supports text plus image or text plus audio. The source says the models are meant for real-time decision making across data centers, RTX workstations, and Jetson edge devices.

    Image from @liquidai's post
  22. Hugging Face BlogOfficialAI score49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    AILiquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  23. Daniel HanXAI score48

    Unsloth enables local training of decision models on 3GB VRAM

    AIUnsloth now lets users train their own decision models locally on just 3GB of VRAM by fine-tuning Qwen, Gemma, and Llama with a Clef head. The post says this raises accuracy from 30% to as high as 78%, and the tool is available through Unsloth Desktop.

    Video from @danielhanchen's post
  24. Aravind SrinivasXAI score62

    Perplexity open-sources pplx-embed-v2-late multimodal embedding models

    AIPerplexity is open-sourcing pplx-embed-v2-late, multi-vector embedding models for text and images in one shared space, in 9B and 0.6B sizes. The 9B model can index multimodal data, the 0.6B model can run queries on device, and PDF pages can be searched without OCR. The author reports 92.4% on MADQA and 64% on BrowseComp+, with weights available on Hugging Face.

    Why it matters: Two open-weight multi-vector models share one space for text and images, with a 0.6B on-device option, a useful comparison for building multimodal retrieval.

  25. Unsloth AIOfficialAI score40

    Unsloth lets users train local decision models on 4GB VRAM

    AIUnsloth released an open-source method to fine-tune LLMs into decision models that run locally, lifting Qwen3.5 0.8B's aggregate accuracy from 20.7% to 74.3% across three decision benchmarks. The team used a Clef head with LoRA (r=64) for one epoch on just 4GB VRAM, with the approach applicable to models such as Qwen3.8 and Gemma 4. A guide and notebooks are available on the Unsloth documentation site and GitHub.

    Image from @UnslothAI's post