Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 5

Oct 5Mon
  1. ThariqXAI score22

    Thariq says HTML planning is more token efficient than raw HTML

    AIThariq says planning with HTML is much more token efficient than generating raw HTML. The model does not need to recreate components or logic for common elements such as state machines, diagrams, and code snippets. Background from the quoted post says he is building a Claude Code skill that generates HTML plans, with linting to reduce common failures.

  2. IThome · AINewsAI score34

    Microsoft Word Copilot adds source citations to curb AI hallucinations

    AIMicrosoft has added citation links to Copilot replies in Word, letting users click through to original web pages or internal documents to verify information. The company says the change improves transparency about where Copilot's information comes from. The feature targets AI hallucinations, which are errors or fabricated sources produced by AI tools.

  3. IThome · AINewsAI score49

    Reflection AI releases open-weight Beam model to rival DeepSeek and Kimi

    AIReflection AI, an Nvidia-backed startup, released Beam, its first open-weight large model, aimed at coding and agent tasks. The company says Beam is comparable to Z.ai's GLM-5.2 and is approaching Qwen3.8-Max on coding and agent work. Beam has 501 billion total parameters, with 23 billion activated per task in a sparse architecture.

  4. SunoOfficialAI score12

    Suno ranks on a16z's Top 100 Gen AI consumer apps list

    AISuno was named to a16z's annual Top 100 Gen AI consumer apps list, placing among only seven companies ranked across its web, mobile, and consumer lists. The company also announced it is hiring across its teams.

  5. Google Developers BlogOfficialAI score62

    EmbeddingGemma 2 releases multimodal embeddings with modular encoder loading

    AIGoogle released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license that maps text, code, images, video, and audio into a shared 768-dimensional space. Developers can load a 270M-parameter text and code setup, or add vision and audio encoders up to a 740M-parameter full multimodal model. Matryoshka truncation to 256 or 128 dimensions reduces vector storage, with the guide noting quality losses on image, video, and speech retrieval at lower dimensions.

    Why it matters: The guide gives concrete encoder sizes and dimension-storage tradeoffs, showing how to choose a configuration for text, code, image, video, and audio retrieval.

  6. Apple Machine Learning ResearchOfficialAI score23

    RISED uses rubrics to guide multi-environment LLM agent training and data selection

    AIApple researchers introduce RISED, a framework that uses rubrics to guide data selection and policy supervision when training one LLM agent across multiple interactive environments. An LLM judge tags rollouts with a shared rubric vocabulary, positive rubrics provide privileged context for an on-policy self-distillation teacher, and negative rubrics steer generation away from recurring failures. The authors report that RISED achieves the highest mean pass rate across environments and ranks first or second in each environment, across model backbones.

  7. Together AI BlogOfficialAI score38

    Together AI Expands Enterprise Inference on IBM Cloud with NVIDIA B300 GPUs

    AITogether AI is running a large dedicated inference cluster of NVIDIA B300 GPUs on IBM Cloud, backed by NVIDIA Spectrum-X Ethernet networking, and is the first customer on it. Together AI operates the inference layer, IBM provides the cloud, and NVIDIA supplies the silicon and networking. The companies say the setup aims to deliver enterprise-grade, open-model inference at scale.

  8. Google Developers BlogOfficialAI score67

    Google releases EmbeddingGemma 2, a multimodal embedding model for on-device search

    AIGoogle DeepMind launched EmbeddingGemma 2, an open-weight 740M parameter model that maps text, images, video frames, and audio into one vector space. The model can run on-device, with about 567MB active RAM for the full multimodal model on a Google Pixel 11 Pro, and is available through Google AI Edge Gallery, Google AI Edge Foresight on Mac, and MediaPipe Tasks, with ML Kit support coming in the weeks ahead.

    Why it matters: The post names concrete on-device apps, memory footprints, and latency figures, showing how a multimodal embedding model can power local search without cloud calls.

  9. Cursor ChangelogOfficialAI score58

    Cursor iOS app adds remote control for local agents on your computer

    AICursor's iOS app now lets users see and reply to local agents running on their computer. Remote control is on by default except for Enterprise organizations, and agents keep running on the computer rather than moving to the cloud. The computer must stay on and online, and users can enable Keep this computer awake in desktop settings.

  10. Tomasz TunguzBlogAI score46

    Vercel Builds an Inbound Sales Agent Run by 14 Rules

    AIVercel's COO Jeanne DeWitt Grosser described how the company built an AI agent that runs the top of its sales funnel, starting from a roughly 125-line prompt written by its best SDR. The team moved the agent from supervised drafting to autonomous operation by August, then split the prompt into 14 deterministic rules and a model-handled judgment layer. Grosser said the system runs inbound for about $1,000 per year in inference and infrastructure.

  11. Goodfire ResearchOfficialAI score62

    Goodfire finds activation probes can detect reward hacking in open-source models

    AIGoodfire Research reports that reward hacking appears in 50–96% of rollouts across three open-source models on three agentic benchmarks. The team found an internal signal tied to cheating and gaming a metric, and simple activation probes catch some hacks that LLM chain-of-thought monitors miss. A probe can screen every transcript cheaply, and in one setup cut LLM monitoring cost by 90% with a roughly 1% precision drop.

    Why it matters: The study links a reward hacking signal in model activations to monitoring cost and detection, showing how probes compare with chain-of-thought monitors on the same runs.

  12. Together AIOfficialAI score46

    Reflection AI launches Beam, a 501B-parameter open agentic model

    AIReflection AI has introduced Beam, an open agentic model with 501B total parameters and 23B active, trained end-to-end from scratch. Full weights are slated for release this month. Together AI congratulated the team and is hosting a NYC meet-up with Reflection and NVIDIA next week.

    Image from @togethercompute's post
  13. Google FlowOfficialAI score13

    Google Flow hosts CultureCon creators on integrating AI tools into their craft

    AIGoogle Flow's CultureCon conversation with Arlander Taylor, Ceej Vega, and Mattaniah Aytenfsu covered how creators intentionally integrate emerging tools into their work. The panelists discussed protecting early brainstorming stages from tools and using AI as a creative partner to overcome blocks. They framed the goal as leveraging these tools to amplify existing skills.

    Image from @FlowbyGoogle's post
  14. Amjad MasadXAI score13

    Replit adds TikTok Ads MCP for promoting apps from its workspace

    AIReplit has made TikTok Ads MCP available inside its workspace, letting developers bring advertising workflows into the same place they build apps. Users connect their TikTok Ads account to run promotion from Replit. The main post, by Amjad Masad, is a short remark that promoting an app on TikTok is the request.

  15. meng shaoXAI score47

    Reflection previews Beam, a 501B-parameter open agentic model

    AIReflection AI previewed Beam, an MoE open model with 501B total and 23B active parameters, claiming 3–4x better inference efficiency than GLM 5.2. The model was pretrained from scratch on 23.8T tokens in four weeks, and its RL run used 10,500 GB300 GPUs over four weeks, which the post describes as possibly the largest publicly recorded. Reflection positions Beam as a workhorse open model for enterprises, governments, and developers, with full weights due this month.

    Image from @shao__meng's post
  16. Aravind SrinivasXAI score35

    Perplexity Mac app adds tabs and multi-window spatial canvas

    AIPerplexity's Mac app now supports tabs, and sessions can open in separate windows for parallel multitasking. Users can run tasks side by side or spread sessions across different screens. The feature is live in version 26.37.1 for all Computer users on Mac.

  17. Ethan MollickXAI score46

    Cowork moves inference and VM to the cloud, with local file access

    AIEthan Mollick reports that he moved much of his complex Cowork work to the new Claude Projects, which persistently chat with a dedicated cloud VM, finding them much better in most ways but poorly documented. Felix Rieseberg, who works on Cowork, explains that the new version runs model inference and the VM in the cloud, with each session in its own sandbox that is destroyed when the session ends. Files are accessed only from folders the user explicitly adds, with the desktop app handling those requests.

  18. Chips and CheeseBlogAI score45

    NVIDIA's Olympus Core Pushes Server Single-Threaded Performance Boundaries

    AINVIDIA's Olympus is a 10-wide out-of-order server core running at 3.3 GHz that prioritizes per-clock performance over high clock speeds. It uses a simultaneous multi-threading (SMT) implementation, unlike Arm's Cortex X925, and has out-of-order structures larger than X925's. In SPEC CPU2026, its branch prediction accuracy is slightly behind AMD's Zen 5 and slightly ahead of Intel's Lion Cove.

  19. Noah ZwebenXAI score40

    Claude can now join Slack group DMs and reply in threads

    AIClaude can be added to Slack group DMs the same way as any other member. It answers in a thread and keeps following that thread, so anyone in the DM can reply to it there. It can also use the personal connectors of whoever asks.

  20. Amjad MasadXAI score8

    Amjad Masad praises family businesses powered by Replit

    AIReplit CEO Amjad Masad expressed pride in the number of family-run businesses powered by Replit. A linked Replit example describes Cockle Finance, where Dan and his father Steve built a custom tool that helped them take on more business and grow 15% in three months.

  21. Dongxi NLPXAI score60

    Reflection AI's Beam open model is compared against leading Chinese models

    AIThe author says Beam, a 501B-parameter open model from Reflection AI, comes close to GLM 5.2 in capability but trails GLM 5.3, Kimi K3, and DeepSeek V4.1 Flash in several areas. The author attributes Beam's competitiveness mainly to inference efficiency, with inference compute at roughly one-third to one-quarter of GLM 5.2's.

  22. Google FlowOfficialAI score10

    Google Flow for Beginners digital workshop starts October 8

    AIGoogle Labs is hosting a beginner-friendly Google Flow digital workshop starting October 8 at 9 AM PT. Led by Google Labs Staff Experience Designer Reed Enger, it offers an in-depth guided walkthrough of the tool. Registration is available through the RSVP link in the post.

    Image from @FlowbyGoogle's post
  23. Sophia YangXAI score62

    Reflection AI's Beam open model has 501B total parameters and 23B active

    AISophia Yang congratulated Reflection AI on Beam, a 501B-parameter open model with 23B active per token. She attributes its efficiency to an RL length penalty that discourages unnecessary tokens and a sparse MoE architecture. Reflection says full weights will be released this month, and the quoted post reports training over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over four weeks.

    Why it matters: The post explains Beam's efficiency through an RL length penalty and sparse MoE design, with benchmark charts comparing it against other open models.

  24. PyTorch BlogOfficialAI score40

    PyTorch consolidates media decoding and encoding in TorchCodec

    AIPyTorch moves all image, video, and audio decoding and encoding into TorchCodec, which handles CPU and CUDA. TorchVision and TorchAudio now focus on transforms, and the older decoding APIs in both libraries are deprecated or removed. All three libraries are now ABI stable, so they no longer need rebuilding for each PyTorch release.

  25. Nous ResearchOfficialAI score18

    Nous Research argues AI agents should give users full control

    AINous Research says users should control their agent's models, data, memory, compute location, prompts, tools, and code. The post lists choices such as switching models mid-conversation, running fully offline, and exporting the agent. It frames these freedoms as the standard an agent should meet, calling it "yours."

  26. Harrison ChaseXAI score50

    Cognition's Devin adds "Dreaming" offline memory cleanup, open-sourced as a standard

    AIHarrison Chase praises Cognition's "Dreaming" feature, which lets Devin clean stale memory records and surface latent information offline. He argues agent memory needs an offline cleanup loop rather than only better retrieval, and questions how inferred memories get validated before use. He also welcomes Cognition's plan to release Agent Memory Repo as an open standard.