Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. Design ArenaOfficialAI score22

    Design Arena says Opus 5.5 builds better websites than Opus 5

    AIDesign Arena reports that websites built by Opus 5.5 show stronger sectioning, visual hierarchy, and balance than those from its predecessor, Opus 5. The post also says Opus 5.5's motion design and animations have improved significantly over Opus 5, which was released just over 2.5 months earlier.

    Video from @DesignArena's post
  2. Design ArenaOfficialAI score44

    Claude Opus 5.5 tops four Design Arena leaderboards after two weeks

    AIAnthropic's Claude Opus 5.5 has taken first place on four Design Arena leaderboards: Overall Frontend, Data Visualization, 3D Design, and React Native. It also ranks in the top three on the Game Dev and UI Components leaderboards, about two weeks after its release. Design Arena says developers, designers, and casual users have embraced the model.

    Image from @DesignArena's post
  3. Marcus on AIBlogAI score62

    Gary Marcus says OpenAI's math result tells us almost nothing yet

    AIGary Marcus says OpenAI's new math result is not the real news, because its report omits the procedure, the model architecture, and the failure rate. He argues that without these details, nobody can tell whether the system generalizes beyond math or only exploits Lean and synthetic data in a verifiable domain. The post includes a companion section by Terence Tao, whose take is not shown in the provided text.

  4. Semafor · TechnologyNewsAI score42

    Higgsfield launches tool for building custom AI influencer characters

    AIHiggsfield has updated its AI influencer tool, letting users insert their own AI-created characters into existing videos. The characters look realistic but deliberately unhuman, with geometric haircuts and elongated necks, as a spokesperson said a flawless face reads as generic AI. The company, which says it has 30 million users, recently announced a $1 billion run rate.

  5. Liquid AIOfficialAI score38

    Liquid AI's Open d1 models run on NVIDIA hardware with llama.cpp support

    AILiquid AI's Open d1 models run across NVIDIA DGX, RTX, and Jetson hardware, with day-one llama.cpp support for deployment anywhere. Measured one request at a time, the d1-3B model's single-question latency is 8 ms on an NVIDIA RTX 4090, 16 ms on Jetson AGX Thor, 26 ms on Jetson AGX Orin 64 GB, and 50 ms on Jetson Orin Nano.

  6. Liquid AIOfficialAI score30

    Liquid AI shows d1-3B running 10 live-camera demos, one forward pass per frame

    AILiquid AI built 10 live-camera demos for its d1-3B model, ranging from gesture-controlled games to content moderation, each using one forward pass per frame. In collaboration with NVIDIA Robotics, the company also showed d1-3B navigating an environment in Isaac Sim, served on a Jetson in a hardware-in-the-loop setup.

    Video from @liquidai's post
  7. Liquid AIOfficialAI score36

    Liquid AI releases d1-omni-600M, a 600M multimodal model for on-device tasks.

    AILiquid AI has released d1-omni-600M, an experimental 600M-parameter model that handles text plus image or audio input. It combines LFM2.5-Encoder-350M with vision and audio encoders and leads the company's text benchmark comparison on toxicity detection and paraphrase identification. The post suggests uses such as voice-command routing, on-device moderation, and intent classification.

    Image from @liquidai's post
  8. Liquid AIOfficialAI score23

    Liquid AI's d1-3B tops sub-10B models on Decision Index v0.2.1

    AILiquid AI's d1-3B ranks first among models under 10B parameters on the Decision Index v0.2.1, a benchmark for structured decision-making. Built from LFM2.5-VL-3B, it makes decisions from text and images in a single pass. It is suited to reranking, agent guardrails, and visual inspection.

    Image from @liquidai's post
  9. Liquid AIOfficialAI score52

    Liquid AI releases open-weight d1-3B and d1-omni-600M multimodal models

    AILiquid AI released Open d1, two open-weight multimodal models in its d1 decision model family. The d1-3B model supports text and vision, while d1-omni-600M supports text plus image or text plus audio. The source says the models are meant for real-time decision making across data centers, RTX workstations, and Jetson edge devices.

    Image from @liquidai's post
  10. Hugging Face BlogOfficialAI score49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    AILiquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  11. Mark ChenXAI score46

    OpenAI's Navier-Stokes progress marks a decade of math advances in a week

    AIMark Chen says the Navier-Stokes achievement matters more for the figure it shows than for the problem itself, representing a decade of mathematical progress in a single week. He says he is eager to apply these tools to life sciences, the building of OpenAI's next models, and alignment research.

    Image from @markchen90's post
  12. GoogleOfficialAI score42

    Google's Project Suncatcher tests TPUs in orbit on a satellite

    AIGoogle launched its first test satellite carrying four TPUs into orbit last week as part of Project Suncatcher, a moonshot exploring whether machine learning infrastructure could one day operate in space. The test aims to determine whether Google's AI hardware can withstand the physical stress of spaceflight and the radiation and thermal extremes of orbit.

    Image from @Google's post
  13. GammaOfficialAI score43

    Gamma 5 rebuilds its engine with an agent, design freedom, and imports

    AIGamma announces Gamma 5, which it calls its biggest update, rebuilding its engine from the ground up. The release adds an agent for brainstorming, research, and editing, plus style control from described looks or visual inspiration. It also supports importing and exporting PowerPoints, PDFs, and company brand, with connections to Slack, Notion, Salesforce, Claude, and ChatGPT.

    Video from @GammaApp's post
  14. Google Cloud TechOfficialAI score34

    Google explains eager vs. lazy loading of MCP tools in Agent Plugins

    AIGoogle DevRel's James O'Reilly explains how Antigravity Agent Plugins expose local MCP server tools to the model, either eagerly as top-level functions or lazily through a call_mcp_tool proxy. Eager loading, set via "eager": true in mcp_config.json, avoids the discovery turn but adds fixed per-turn token overhead that can degrade reasoning with 100+ tools. Lazy loading is the plugin default and keeps baseline token use low, at the cost of an extra proxy hop and a higher chance of JSON quoting errors.

  15. Databricks BlogOfficialAI score41

    Databricks Apps Adds On-Behalf-of-User Authorization for Permission-Aware Apps

    AIDatabricks announced general availability of on-behalf-of-user (OBO) authorization for Databricks Apps, letting apps act with the signed-in user's identity so Unity Catalog enforces that user's row filters and column masks. Developers can request narrow API scopes such as sql:restricted-query, which allows only read-only SQL queries, while apps keep a dedicated service principal for app-owned operations.

  16. Aravind SrinivasXAI score62

    Perplexity open-sources pplx-embed-v2-late multimodal embedding models

    AIPerplexity is open-sourcing pplx-embed-v2-late, multi-vector embedding models for text and images in one shared space, in 9B and 0.6B sizes. The 9B model can index multimodal data, the 0.6B model can run queries on device, and PDF pages can be searched without OCR. The author reports 92.4% on MADQA and 64% on BrowseComp+, with weights available on Hugging Face.

    Why it matters: Two open-weight multi-vector models share one space for text and images, with a 0.6B on-device option, a useful comparison for building multimodal retrieval.

  17. The Robot ReportNewsAI score26

    Teradyne Robotics and Elite Robots settle cobot software dispute

    AITeradyne Robotics and Elite Robots reached a mutual agreement ending a legal dispute over robots and software, announced Oct. 1 without disclosing settlement terms. The case followed a preliminary injunction issued by the Regional Court of Hamburg against Elite Robots Deutschland over allegedly infringing Universal Robots software, which was not a final finding of infringement. Elite Robots did not admit liability.

  18. DeedyXAI score46

    OpenAI's math results spark claims of AGI and Millennium Prize progress

    AIDeedy argues LLMs have made substantial progress on four of the seven Millennium Prize problems, including a claimed Navier-Stokes result, conditional on verification. He says OpenAI's results averaged only 3 hours of thinking compute on unreleased models. He concludes that by most definitions of AGI, we have already achieved it.

  19. GitHub Copilot ChangelogOfficialAI score30

    GitHub launches purpose-built AI model for leaked secret detection across developer workflows

    AIGitHub is rolling out a fine-tuned, purpose-built model for secret detection that reads surrounding code to identify likely credentials, including passwords without recognizable token formats. Existing AI-detected Password alerts have been upgraded automatically, and AI-detected secrets in push protection is in private preview. New opt-in checks in push protection and the GitHub Copilot /security-review command will consume GitHub AI Credits.

  20. PerplexityOfficialAI score32

    Perplexity reports 92.4% on MADQA document QA benchmark

    AIPerplexity reports its system reached 92.4% accuracy on MADQA, a benchmark of 500 questions over 800 PDFs. The post describes this as the top result among retrievers on that benchmark.

    Image from @perplexity_ai's post
  21. PerplexityOfficialAI score41

    Perplexity's 0.6B and 9B embedding models share one embedding space

    AIPerplexity's 0.6B and 9B models are both distilled token by token from one 18B teacher, so they share a single embedding space. A corpus indexed with the 9B model can be searched using 0.6B queries, raising ViDoRe v3 from 62.3% to 63.5% with no added query cost.

    Image from @perplexity_ai's post
  22. PerplexityOfficialAI score36

    Perplexity's models search PDFs and slides directly, without OCR

    AIPerplexity's models embed images and rendered pages directly, so PDFs, slides, and scans can be searched without OCR. This preserves tables, figures, and layout that text extraction typically drops.

    Image from @perplexity_ai's post
  23. PerplexityOfficialAI score36

    Perplexity's pplx-embed-v2-late keeps per-token vectors for retrieval

    AIPerplexity's pplx-embed-v2-late retains a 128-dimensional vector for each token rather than compressing a document into one vector. It scores matches with MaxSim, pairing each query token with its closest document token, which the post presents as preserving detail in long or visually dense pages.

    Image from @perplexity_ai's post
  24. Microsoft ResearchOfficialAI score24

    Agent Lightning connects existing AI agents to reinforcement learning training

    AIMicrosoft Research introduced Agent Lightning, a tool that connects existing AI agents to reinforcement learning training. It aims to make agents easier to improve without rebuilding them, since their tools, context, and decision-making are typically managed by complex frameworks.

    Video from @MSFTResearch's post
  25. Codex · GitHub ReleasesOfficialAI score34

    Codex 0.161.0 makes GPT-6.1 Sol default and adds Daybreak opt-in

    AIOpenAI's Codex 0.161.0 release makes GPT-6.1 Sol the default model in the bundled and Amazon Bedrock catalogs. It adds Amazon Bedrock support for multi-agent V2 and Ultra reasoning on compatible models, plus opt-in Daybreak routing enabled through --enable cli_daybreak or features.cli_daybreak=true.

  26. Microsoft ResearchOfficialAI score62

    Microsoft Research Asia releases Agent Lightning v1.0 for agentic RL with real harnesses

    AIMicrosoft Research Asia has open-sourced Agent Lightning v1.0, a roughly 3,500-line agentic RL framework that trains the same agent harness used in deployment. In an end-to-end coding agent pipeline, Qwen3.5-9B rose from 41.8% to 56.4% Pass@1 on SWE-bench Verified using about 6,000 training samples. The framework runs agents as standard Kubernetes jobs without paid commercial sandbox services.

    Why it matters: The source shows how training with the deployed agent harness avoids rebuilding agents, and reports concrete SWE-bench Verified gains from about 6,000 samples.

  27. NVIDIA Technical BlogOfficialAI score22

    Validate AI Factory Changes with Digital Twins and AI Agents

    AINVIDIA describes using digital twins and AI agents to validate changes to AI factory infrastructure, which combines GPUs, CPUs, switches, DPUs, and SuperNICs with schedulers, orchestration services, security controls, and a fast-changing software stack. The source frames the challenge as confirming that hardware, software, and policies work together for target workloads before deployment. The available excerpt does not give further detail on specific tools or results.