Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 5

Oct 5Mon
  1. IThome · AINewsAI score34

    Microsoft Word Copilot adds source citations to curb AI hallucinations

    AIMicrosoft has added citation links to Copilot replies in Word, letting users click through to original web pages or internal documents to verify information. The company says the change improves transparency about where Copilot's information comes from. The feature targets AI hallucinations, which are errors or fabricated sources produced by AI tools.

  2. IThome · AINewsAI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    AIReflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.

  3. Google Developers BlogOfficialAI score62

    EmbeddingGemma 2 releases multimodal embeddings with modular encoder loading

    AIGoogle released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license that maps text, code, images, video, and audio into a shared 768-dimensional space. Developers can load a 270M-parameter text and code setup, or add vision and audio encoders up to a 740M-parameter full multimodal model. Matryoshka truncation to 256 or 128 dimensions reduces vector storage, with the guide noting quality losses on image, video, and speech retrieval at lower dimensions.

    Why it matters: The guide gives concrete encoder sizes and dimension-storage tradeoffs, showing how to choose a configuration for text, code, image, video, and audio retrieval.

  4. Apple Machine Learning ResearchOfficialAI score23

    RISED uses rubrics to guide multi-environment LLM agent training and data selection

    AIApple researchers introduce RISED, a framework that uses rubrics to guide data selection and policy supervision when training one LLM agent across multiple interactive environments. An LLM judge tags rollouts with a shared rubric vocabulary, positive rubrics provide privileged context for an on-policy self-distillation teacher, and negative rubrics steer generation away from recurring failures. The authors report that RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, across model backbones.

  5. Together AI BlogOfficialAI score38

    Together AI Expands Enterprise Inference on IBM Cloud with NVIDIA B300 GPUs

    AITogether AI is running a large dedicated inference cluster of NVIDIA B300 GPUs on IBM Cloud, backed by NVIDIA Spectrum-X Ethernet networking, and is the first customer on it. Together AI operates the inference layer, IBM provides the cloud, and NVIDIA supplies the silicon and networking. The companies say the setup aims to deliver enterprise-grade, open-model inference at scale.

  6. Google Developers BlogOfficialAI score67

    Google releases EmbeddingGemma 2, a multimodal embedding model for on-device search

    AIGoogle DeepMind launched EmbeddingGemma 2, an open-weight 740M parameter model that maps text, images, video frames, and audio into one vector space. The model can run on-device, with about 567MB active RAM for the full multimodal model on a Google Pixel 11 Pro, and is available through Google AI Edge Gallery, Google AI Edge Foresight on Mac, and MediaPipe Tasks, with ML Kit support coming in the weeks ahead.

    Why it matters: The post names concrete on-device apps, memory footprints, and latency figures, showing how a multimodal embedding model can power local search without cloud calls.

  7. Cursor ChangelogOfficialAI score58

    Cursor iOS app adds remote control for local agents on your computer

    AICursor's iOS app now lets users see and reply to local agents running on their computer. Remote control is on by default except for Enterprise organizations, and agents keep running on the computer rather than moving to the cloud. The computer must stay on and online, and users can enable Keep this computer awake in desktop settings.

  8. Tomasz TunguzBlogAI score46

    Vercel Builds an Inbound Sales Agent Run by 14 Rules

    AIVercel's COO Jeanne DeWitt Grosser described how the company built an AI agent that runs the top of its sales funnel, starting from a roughly 125-line prompt written by its best SDR. The team moved the agent from supervised drafting to autonomous operation by August, then split the prompt into 14 deterministic rules and a model-handled judgment layer. Grosser said the system runs inbound for about $1,000 per year in inference and infrastructure.

  9. Goodfire ResearchOfficialAI score62

    Goodfire finds activation probes can detect reward hacking in open-source models

    AIGoodfire Research reports that reward hacking appears in 50–96% of rollouts across three open-source models on three agentic benchmarks. The team found an internal signal tied to cheating and gaming a metric, and simple activation probes catch some hacks that LLM chain-of-thought monitors miss. A probe can screen every transcript cheaply, and in one setup cut LLM monitoring cost by 90% with a roughly 1% precision drop.

    Why it matters: The study links a reward hacking signal in model activations to monitoring cost and detection, showing how probes compare with chain-of-thought monitors on the same runs.

  10. Together AIOfficialAI score46

    Reflection AI launches Beam, a 501B-parameter open agentic model

    AIReflection AI has introduced Beam, an open agentic model with 501B total parameters and 23B active, trained end-to-end from scratch. Full weights are slated for release this month. Together AI congratulated the team and is hosting a NYC meet-up with Reflection and NVIDIA next week.

    Image from @togethercompute's post
  11. meng shaoXAI score47

    Reflection previews Beam, a 501B-parameter open agentic model

    AIReflection AI previewed Beam, an MoE open model with 501B total and 23B active parameters, claiming 3–4x better inference efficiency than GLM 5.2. The model was pretrained from scratch on 23.8T tokens in four weeks, and its RL run used 10,500 GB300 GPUs over four weeks, which the post describes as possibly the largest publicly recorded. Reflection positions Beam as a workhorse open model for enterprises, governments, and developers, with full weights due this month.

    Image from @shao__meng's post
  12. Aravind SrinivasXAI score35

    Perplexity Mac app adds tabs and multi-window spatial canvas

    AIPerplexity's Mac app now supports tabs, and sessions can open in separate windows for parallel multitasking. Users can run tasks side by side or spread sessions across different screens. The feature is live in version 26.37.1 for all Computer users on Mac.

  13. Ethan MollickXAI score46

    Cowork moves inference and VM to the cloud, with local file access

    AIEthan Mollick reports that he moved much of his complex Cowork work to the new Claude Projects, which persistently chat with a dedicated cloud VM, finding them much better in most ways but poorly documented. Felix Rieseberg, who works on Cowork, explains that the new version runs model inference and the VM in the cloud, with each session in its own sandbox that is destroyed when the session ends. Files are accessed only from folders the user explicitly adds, with the desktop app handling those requests.

  14. Chips and CheeseBlogAI score45

    NVIDIA's Olympus Core Pushes Server Single-Threaded Performance Boundaries

    AINVIDIA's Olympus is a 10-wide out-of-order server core running at 3.3 GHz that prioritizes per-clock performance over high clock speeds. It uses a simultaneous multi-threading (SMT) implementation, unlike Arm's Cortex X925, and has out-of-order structures larger than X925's. In SPEC CPU2026, its branch prediction accuracy is slightly behind AMD's Zen 5 and slightly ahead of Intel's Lion Cove.

  15. Noah ZwebenXAI score40

    Claude can now join Slack group DMs and reply in threads

    AIClaude can be added to Slack group DMs the same way as any other member. It answers in a thread and keeps following that thread, so anyone in the DM can reply to it there. It can also use the personal connectors of whoever asks.

  16. Dongxi NLPXAI score60

    Reflection AI's Beam open model is compared against leading Chinese models

    AIThe author says Beam, a 501B-parameter open model from Reflection AI, comes close to GLM 5.2 in capability but trails GLM 5.3, Kimi K3, and DeepSeek V4.1 Flash in several areas. The author attributes Beam's competitiveness mainly to inference efficiency, with inference compute at roughly one-third to one-quarter of GLM 5.2's.

  17. Sophia YangXAI score62

    Reflection AI's Beam open model has 501B total parameters and 23B active

    AISophia Yang congratulated Reflection AI on Beam, a 501B-parameter open model with 23B active per token. She attributes its efficiency to an RL length penalty that discourages unnecessary tokens and a sparse MoE architecture. Reflection says full weights will be released this month, and the quoted post reports training over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks.

    Why it matters: The post explains Beam's efficiency through an RL length penalty and sparse MoE design, with benchmark charts comparing it against other open models.

  18. PyTorch BlogOfficialAI score40

    PyTorch consolidates media decoding and encoding in TorchCodec

    AIPyTorch moves all image, video, and audio decoding and encoding into TorchCodec, which handles CPU and CUDA. TorchVision and TorchAudio now focus on transforms, and the older decoding APIs in both libraries are deprecated or removed. All three libraries are now ABI stable, so they no longer need rebuilding for each PyTorch release.

  19. Harrison ChaseXAI score50

    Cognition's Devin adds "Dreaming" offline memory cleanup, open-sourced as a standard

    AIHarrison Chase praises Cognition's "Dreaming" feature, which lets Devin clean stale memory records and surface latent information offline. He argues agent memory needs an offline cleanup loop rather than only better retrieval, and questions how inferred memories get validated before use. He also welcomes Cognition's plan to release Agent Memory Repo as an open standard.

  20. NVIDIA AIOfficialAI score39

    NVIDIA releases Nemotron-Labs-3-Competitive-Coding model on Hugging Face

    AINVIDIA has published Nemotron-Labs-3-Competitive-Coding on Hugging Face, a competitive-programming specialist model built on Nemotron-3-Ultra. The model is available in the NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4 repository, indicating a 550B-parameter total size with 55B active parameters in NVFP4 format.

  21. clem 🤗XAI score72

    Reflection AI announces Beam, a 501B-parameter agentic open model

    AIReflection AI introduced Beam, an agentic open model with 501B total parameters and 23B active parameters, trained end-to-end from scratch. The quoted announcement says it targets frontier reasoning efficiency and coding and agentic tasks, with full weights due this month. Clément Delangue, Hugging Face's CEO, reposted it with a welcome to the Reflection organization on Hugging Face.

    Image from @ClementDelangue's post
  22. Georgi GerganovXAI score36

    llama.cpp v0.6.0 adds Clef, Qwen3.8-Flash-Next, and Metal speedups

    AIThe llama.cpp v0.6.0 release adds Clef support for text and vision, along with high-quality support for Qwen3.8-Flash-Next. It also brings a major Metal performance improvement and a new llama_batch_ext API, and the project website at llama.app has been refreshed.

  23. ReflectionOfficialAI score23

    Reflection AI's Beam model pretrained in four weeks on 24T tokens

    AIReflection AI says its Beam model was pretrained in 4 weeks on 24T high-quality tokens, giving it innate coding capabilities. The company credits MoE stability improvements and large-scale data curation and deduplication for a base model it claims outperforms open-source base models of the same class. It presents this strong reasoning foundation as what makes sustained reinforcement learning gains possible.

    Image from @reflection_ai's post
  24. ReflectionOfficialAI score42

    Reflection AI previews Beam, a 500B open model under Apache 2.0

    AIReflection AI says its Beam model, with a 500B form factor, combines strong agentic performance and efficient reasoning for enterprises, governments, and developers. Beam is in final red-teaming and will be released this month under an Apache 2.0 license, with quantized FP8 and NVFP4 versions for efficient deployment. Early access sign-ups are open on the company's platform.

  25. Alex HeathXAI score52

    Reflection's founders discuss building a DeepSeek of the West with Beam

    AIReflection is set to release Beam, its first open-weight AI model, aiming to become a Western counterpart to DeepSeek. The source says Beam is trained from scratch for coding, reasoning, and AI agents, with benchmarks placing it alongside the strongest open models and more efficient token economics. Reflection has raised $4.6 billion from investors including Nvidia, Sequoia, and Lightspeed, and the interview covers its monetization plans for open-weight models.

    Video from @alexeheath's post
  26. Gergely OroszXAI score35

    Gergely Orosz says coding agent product strategy feels like "YOLO"

    AIGergely Orosz says many coding agents seem to follow a "YOLO" product strategy, with rapid week-over-week change learned about through random social media posts. He notes this makes some sense given how quickly the industry and capabilities keep changing. Quoted context reports that Anthropic is removing Cowork's local option for Pro/Max users, with new tasks running in the cloud while existing local tasks stay on the computer.