Skip to contentSkip to stories

Updated

Open source

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. IThome · AINewsAI score41

    Strata engine runs 125B Qwen3.8 model on 12GB GPU at 94 tokens/s

    AIDeveloper Niko1221 has open-sourced Strata, an engine that runs a quantized 125B-parameter Qwen3.8-Flash-Next model on consumer GPUs with at least 12GB of VRAM. Strata loads the MoE model into RAM and keeps only frequently used experts in VRAM, and uses a lightweight model for speculative decoding. On an NVIDIA RTX 5070 with 12GB VRAM, the Q2_0 quantization reaches 94 tokens per second for output.

  2. Latent SpaceBlogAI score60

    Reflection launches Beam, a 501B-parameter open-weight coding model

    AIReflection announced Beam, a text-only 501B-total, 23B-active MoE model for coding, agentic, and scientific work, trained from scratch with full weights under Apache 2.0 promised this month. Self-reported results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2, while the roundup notes that GLM 5.3, Kimi K3, Qwen 3.8 Max, and DeepSeek V4.1 Flash are generally ahead.

  3. Claude BlogOfficialAI score62

    Claude now works inside Google Docs, Sheets, and Slides in public beta

    AIClaude for Google Workspace is in public beta on all paid Claude plans, adding a sidebar to Google Docs, Sheets, and Slides. It can read the open file, edit text, build formulas, pivot tables, charts, and slides, and it asks for approval before changes unless the user chooses "Accept all edits." New Docs, Sheets, and Slides connectors in beta let Claude create and edit Google files from the chat, with access matching existing Google sharing permissions.

    Why it matters: The source specifies how Claude edits Docs, Sheets, and Slides in place and where users keep control, which clarifies the practical workflow change.

  4. Artificial Analysis ArticlesOfficialAI score54

    Mistral Large 4 Preview scores 38 on Artificial Analysis Intelligence Index

    AIMistral has released Mistral Large 4 in Research Public Preview, with open weights for the 1T parameter (49B active) model planned for the end of October. It scores 38 on the Artificial Analysis Intelligence Index, comparable to GPT-6 Luna (max, 38) and DeepSeek V4.1 Flash (max, 39), and 50 on the Cyber Index. The source calls it the most intelligent model from outside the US and China, and notes costs of $1.13 per Intelligence Index task at standard pricing.

  5. METR BlogOfficialAI score31

    AI Agents Could Hide Misbehavior by Exploiting Inspect Transcript Viewer

    AIMETR tested whether an AI agent running in an Inspect evaluation could alter the transcript humans review, and a researcher found a vulnerability in about 10 minutes that allowed arbitrary changes to what the reviewer sees. The exploit affects only the displayed transcript, not the underlying data stored in METR's database, and METR has not observed agents using it in its evaluations. METR argues that AI outputs such as transcripts and reasoning should be treated as untrusted input, with monitoring systems treated as security-critical infrastructure.

Oct 5

Oct 5Mon
  1. Teknium 🪽XAI score23

    Teknium publishes catalog of open hardware for Hermes Agent

    AITeknium has created a catalog of open platform hardware and devices that Hermes Agent, or any agent, can build on, integrate with, or run inside. The post is a brief announcement that links to the catalog at with no further specifications or pricing given.

    Image from @Teknium's post
  2. IThome · AINewsAI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    AIReflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.

  3. Google Developers BlogOfficialAI score62

    EmbeddingGemma 2 releases multimodal embeddings with modular encoder loading

    AIGoogle released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license that maps text, code, images, video, and audio into a shared 768-dimensional space. Developers can load a 270M-parameter text and code setup, or add vision and audio encoders up to a 740M-parameter full multimodal model. Matryoshka truncation to 256 or 128 dimensions reduces vector storage, with the guide noting quality losses on image, video, and speech retrieval at lower dimensions.

    Why it matters: The guide gives concrete encoder sizes and dimension-storage tradeoffs, showing how to choose a configuration for text, code, image, video, and audio retrieval.

  4. Google Developers BlogOfficialAI score67

    Google releases EmbeddingGemma 2, a multimodal embedding model for on-device search

    AIGoogle DeepMind launched EmbeddingGemma 2, an open-weight 740M parameter model that maps text, images, video frames, and audio into one vector space. The model can run on-device, with about 567MB active RAM for the full multimodal model on a Google Pixel 11 Pro, and is available through Google AI Edge Gallery, Google AI Edge Foresight on Mac, and MediaPipe Tasks, with ML Kit support coming in the weeks ahead.

    Why it matters: The post names concrete on-device apps, memory footprints, and latency figures, showing how a multimodal embedding model can power local search without cloud calls.

  5. Together AIOfficialAI score46

    Reflection AI launches Beam, a 501B-parameter open agentic model

    AIReflection AI has introduced Beam, an open agentic model with 501B total parameters and 23B active, trained end-to-end from scratch. Full weights are slated for release this month. Together AI congratulated the team and is hosting a NYC meet-up with Reflection and NVIDIA next week.

    Image from @togethercompute's post
  6. meng shaoXAI score47

    Reflection previews Beam, a 501B-parameter open agentic model

    AIReflection AI previewed Beam, an MoE open model with 501B total and 23B active parameters, claiming 3–4x better inference efficiency than GLM 5.2. The model was pretrained from scratch on 23.8T tokens in four weeks, and its RL run used 10,500 GB300 GPUs over four weeks, which the post describes as possibly the largest publicly recorded. Reflection positions Beam as a workhorse open model for enterprises, governments, and developers, with full weights due this month.

    Image from @shao__meng's post
  7. RadixArkOfficialAI score14

    RadixArk joins three SF Tech Week events on open-source AI and RL

    AIRadixArk is speaking at three SF Tech Week events this week on open-source AI, reinforcement learning, and AI infrastructure. Mao Cheng joins an October 7 panel with Novita Labs on the latest Miles post-training release, inference, and agent guardrails. Shi Dong will share research on Miles and recursive self-improvement at an October 8 MiniMax event, and SGLang core contributor Yuwei An will speak at an October 8 AI infra meetup with SkyPilot and H Company.

    Image from @radixark's post
  8. Dongxi NLPXAI score60

    Reflection AI's Beam open model is compared against leading Chinese models

    AIThe author says Beam, a 501B-parameter open model from Reflection AI, comes close to GLM 5.2 in capability but trails GLM 5.3, Kimi K3, and DeepSeek V4.1 Flash in several areas. The author attributes Beam's competitiveness mainly to inference efficiency, with inference compute at roughly one-third to one-quarter of GLM 5.2's.

  9. Sophia YangXAI score62

    Reflection AI's Beam open model has 501B total parameters and 23B active

    AISophia Yang congratulated Reflection AI on Beam, a 501B-parameter open model with 23B active per token. She attributes its efficiency to an RL length penalty that discourages unnecessary tokens and a sparse MoE architecture. Reflection says full weights will be released this month, and the quoted post reports training over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks.

    Why it matters: The post explains Beam's efficiency through an RL length penalty and sparse MoE design, with benchmark charts comparing it against other open models.

  10. PyTorch BlogOfficialAI score40

    PyTorch consolidates media decoding and encoding in TorchCodec

    AIPyTorch moves all image, video, and audio decoding and encoding into TorchCodec, which handles CPU and CUDA. TorchVision and TorchAudio now focus on transforms, and the older decoding APIs in both libraries are deprecated or removed. All three libraries are now ABI stable, so they no longer need rebuilding for each PyTorch release.

  11. Nous ResearchOfficialAI score18

    Nous Research argues AI agents should give users full control

    AINous Research says users should control their agent's models, data, memory, compute location, prompts, tools, and code. The post lists choices such as switching models mid-conversation, running fully offline, and exporting the agent. It frames these freedoms as the standard an agent should meet, calling it "yours."

  12. NVIDIA AIOfficialAI score39

    NVIDIA releases Nemotron-Labs-3-Competitive-Coding model on Hugging Face

    AINVIDIA has published Nemotron-Labs-3-Competitive-Coding on Hugging Face, a competitive-programming specialist model built on Nemotron-3-Ultra. The model is available in the NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 repository, indicating a 550B-parameter total size with 55B active parameters in NVFP4 format.

  13. clem 🤗XAI score72

    Reflection AI announces Beam, a 501B-parameter agentic open model

    AIReflection AI introduced Beam, an agentic open model with 501B total parameters and 23B active parameters, trained end-to-end from scratch. The quoted announcement says it targets frontier reasoning efficiency and coding and agentic tasks, with full weights due this month. Clément Delangue, Hugging Face's CEO, reposted it with a welcome to the Reflection organization on Hugging Face.

    Why it matters: The quoted announcement names Beam's parameter scale, active-parameter count, and coding and agentic focus, which helps readers gauge where it fits among open models.

    Image from @ClementDelangue's post
  14. Georgi GerganovXAI score36

    llama.cpp v0.6.0 adds Clef, Qwen3.8-Flash-Next, and Metal speedups

    AIThe llama.cpp v0.6.0 release adds Clef support for text and vision, along with high-quality support for Qwen3.8-Flash-Next. It also brings a major Metal performance improvement and a new llama_batch_ext API, and the project website at llama.app has been refreshed.

  15. ReflectionOfficialAI score42

    Reflection AI previews Beam, a 500B open model under Apache 2.0

    AIReflection AI says its Beam model, with a 500B form factor, combines strong agentic performance and efficient reasoning for enterprises, governments, and developers. Beam is in final red-teaming and will be released this month under an Apache 2.0 license, with quantized FP8 and NVFP4 versions for efficient deployment. Early access sign-ups are open on the company's platform.

  16. ReflectionOfficialAI score14

    Reflection AI says Beam leads in inference efficiency

    AIReflection AI says its model Beam is 3-4x more efficient than GLM 5.2 and more than 4x more efficient than leading Western open models in inference. The company says this means Beam completes tasks faster and cheaper.

    Image from @reflection_ai's post
  17. Alex HeathXAI score52

    Reflection's founders discuss building a DeepSeek of the West with Beam

    AIReflection is set to release Beam, its first open-weight AI model, aiming to become a Western counterpart to DeepSeek. The source says Beam is trained from scratch for coding, reasoning, and AI agents, with benchmarks placing it alongside the strongest open models and more efficient token economics. Reflection has raised $4.6 billion from investors including Nvidia, Sequoia, and Lightspeed, and the interview covers its monetization plans for open-weight models.

    Video from @alexeheath's post
  18. CognitionOfficialAI score58

    Cognition's Devin adds Dreaming, a nightly memory graph across sessions

    AICognition introduces Dreaming, a feature in which Devin builds a memory graph of how a user likes to work across sessions. At night, Devin self-improves this memory by removing stale records and discovering latent information. Cognition also says it is creating an open-source standard called Agent Memory Repo, linked in the post.

    Video from @cognition's post
  19. Liquid AI · new models on Hugging FaceOfficialAI score44

    LiquidAI releases d1-omni-600M, a 600M decision model for text, image and audio

    AILiquidAI has released d1-omni-600M on Hugging Face, a 587M-parameter model that answers named yes/no, choice and score questions over text, images or up to 30 seconds of speech in a single forward pass. It returns typed answers with zero output tokens by reading the model's distribution over options, and is built on LFM2.5-Encoder-350M with a 16,384-token context length. The model is not a chat model and does not generate text.