Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 25

Sep 25Fri
  1. Stephanie PalazzoloXAI score32

    Fal and Fireworks weigh funding rounds at up to $30B valuation

    AIInference providers Fal and Fireworks are reportedly considering new funding rounds amid soaring inference demand. Fal has discussed raising at a $15B valuation, while Fireworks has considered a $30B valuation, nearly doubling their valuations from earlier rounds this year.

  2. Google AntigravityOfficialAI score34

    Antigravity 2.0 adds planning mode with /plan command

    AIGoogle Antigravity 2.0 now includes a dedicated planning mode, matching the Antigravity CLI. Typing /plan makes the agent research the task and generate an implementation plan for user review before execution, requiring approval to proceed. Users can also request a lighter plan through a natural prompt.

    Video from @antigravity's post
  3. LMSYS OrgOfficialAI score38

    SGLang adds multi-item scoring for faster decision model serving

    AISGLang's /v1/score endpoint returns scores for exact requested labels such as Yes/No or A/B/C, and its multi-item scoring (MIS) computes shared context once while keeping candidates isolated. On Qwen3-8B, 16-candidate p95 latency dropped from 54.1 ms with Generate to 20.6 ms with MIS. On Qwen3-0.6B, MIS p95 stayed under about 100 ms as load rose, versus seconds for Generate and SIS.

    Image from @lmsysorg's post
  4. Kevin Weil 🇺🇸XAI score75

    Claude solves nine-loop scattering amplitude calculation past prior eight-loop record

    AIAnthropic reports that Claude solved a nine-loop calculation in the planar N=4 super-Yang-Mills model, surpassing the previous eight-loop record set by Lance Dixon and collaborators. The quoted post says Claude ran largely unsupervised for days in Claude Science using a single prompt, at a total cost of a few thousand dollars, and Dixon independently verified the result. Kevin Weil's own text praises the achievement and expects AI to advance high energy physics over the coming 12 months.

    Why it matters: The quoted Anthropic post gives a concrete benchmark: Claude ran for days to reach nine loops, extending the previous eight-loop record in a physics model.

  5. VercelOfficialAI score31

    Klaviyo Ships 356 Internal Apps in Two Weeks on Vercel

    AIKlaviyo built an internal app platform on Vercel, and in the first two weeks 512 employees shipped 356 projects. Teams can go from idea to a live app in about three minutes, with full-stack apps running on Klaviyo's databases. Deployments are SSO-gated and private by default.

  6. Boris ChernyXAI score62

    Claude Tag in Slack gains personal connectors for channel workflows

    AIClaude Tag in Slack can now use users' personal connectors, such as Drive, Salesforce, and warehouse access, within channels. The author says Tag writes over 50% of their PRs daily and handles nearly all their data analysis and many product bug fixes. The personal connector feature is available on Teams today and Enterprise next week.

  7. Noah ZwebenXAI score46

    Anthropic shows Claude Code /remote-control demo with Opus 5.5 claymation video

    AIAnthropic's Noah Zweben shared a claymation video showing Claude Code's /remote-control feature, now made with Opus 5.5 after an earlier Opus 4.6 version. The feature, which lets users control Claude Code remotely, is rolling out to Pro users at 10% and ramping, with Team and Enterprise support coming later.

    Video from @noahzweben's post
  8. Google Cloud · AI & Machine LearningOfficialAI score43

    Google Cloud Introduces Managed Reinforcement Learning Fine-Tuning for Gemini Models

    AIGoogle Cloud has launched a managed reinforcement learning fine-tuning service (RLFT) that lets customers adapt Gemini models using a reward function they define instead of labeled answers. Users supply prompts and a reward function, while Google handles the RL infrastructure and proprietary model internals. The guide advises exhausting prompting and supervised fine-tuning first, and notes that RLFT suits tasks that are easy to score but hard to demonstrate.

  9. SemiAnalysisBlogAI score59

    China Holds Over 24GW of Datacenter Capacity, Shifting Inland With AI Demand

    AISemiAnalysis's China Datacenter Model tracks over 1,000 facilities across 60+ operators and puts China's fleet above 24GW, larger than EMEA. The report attributes the buildout to the Eastern Data, Western Compute policy and hyperscale AI demand, which is moving capacity to western hubs such as Inner Mongolia at construction speeds it says the West cannot match.

  10. OpenCodeOfficialAI score22

    OpenCode makes $60 of DeepSeek v4.1 Flash usage permanent

    AIOpenCode says its $60 of usage for DeepSeek v4.1 Flash is now permanent, as part of its "Operation Cheepseek Phase 2" promotion. The post gives no further details on terms, duration, or eligibility.

  11. Amazon ScienceOfficialAI score38

    Amazon and Reactor build kernel path to real-time video generation on Trainium

    AIUsing the Neuron Kernel Interface, Reactor and Amazon's Neuron Science team built a kernel-centric path to real-time autoregressive diffusion video generation on Trainium. They addressed dynamic shapes, memory access patterns, and cache management, which are hard for generic compilers, and developed techniques intended to generalize across models.

  12. Meituan LongCatOfficialAI score62

    Meituan LongCat-2.5-Preview Launches with 1.6T Parameters and 1M-Token Context

    AIMeituan's LongCat team has released LongCat-2.5-Preview, a natively multimodal model with 1.6T total parameters, about 48B active, and a 1M-token context window. The model is built for long-horizon tasks spanning terminals, browsers, GUIs, spreadsheets, and design tools. It is available now through an API on the LongCat platform and a chat interface.

    Image from @Meituan_LongCat's post
  13. Microsoft CopilotOfficialAI score40

    Microsoft Copilot app refreshed to unify chat, agents, app building, and workflows

    AIMicrosoft has refreshed its Copilot app to bring chat, task delegation, app building, and workflow automation into one place. The update is positioned as an AI built for work, with Satya Nadella describing Copilot as a new OS for work spanning models, form factors, and tasks. The announcement includes Autopilot, an enterprise agent, Code for building apps hosted within a company's tenant, Home combining Chat and Cowork, and Office fully embedded in Copilot.

  14. Satya NadellaXAI score52

    Satya Nadella announces Copilot update with Autopilot, Code, Home, and Office

    AIMicrosoft CEO Satya Nadella announced what he called the biggest Copilot update to date, positioning Copilot as a new operating system for work. The update bundles Autopilot, a proactive long-running enterprise agent; Code, for building apps hosted inside a company's tenant; Home, combining Chat and Cowork; and Office, now fully embedded in Copilot. Copilot can also be invoked in Teams, and a new proactive experience called Today surfaces key information from across M365 without a prompt.

    Video from @satyanadella's post

Sep 24

Sep 24Thu
  1. ModelScopeOfficialAI score23

    NeoHorse-Jev-4B open model turns app states into structured decisions

    AIModelScope has released NeoHorse-Jev-4B, a compact open model that converts application states into structured decisions and probabilities. It scores 77.70 across six text decision benchmark groups, ranking first among four open-weight models with complete results in the comparison. Its prefill-only inference supports Choice, Noul, and Score primitives, accepts text or a single image with text, and is available under Apache 2.0 for deployment via vLLM, SGLang, Python, CLI, or HTTP.

    Video from @ModelScope2022's post
  2. vLLMOfficialAI score30

    vLLM integrates TileRT with PD disaggregation, benchmark config published

    AIvLLM has published a blog post explaining how its integration with TileRT works, alongside a public benchmark configuration in the InferenceX repository. The benchmark script covers GLM-5.3 FP8 on MI355X hardware and is linked on GitHub. The post itself provides the integration details.

  3. vLLMOfficialAI score42

    TileRT and vLLM hit 469 tok/s on GLM-5.3 with MI355X

    AIThe TileRT and AMD teams reached 469 tok/s single-user decode for GLM-5.3 on 8× MI355X using vLLM. The setup disaggregates work, with vLLM handling prefill and TileRT handling latency-critical decode through vLLM's V1 connector interface. SemiAnalysis's AgentX benchmark reports the configuration at 470 TPS on GLM 5.3 (FP8), over 40% faster than GB300 TRTLLM using FP4.

  4. Redwood Research BlogBlogAI score41

    Continual learning could make AI monitors that block actions nearly useless

    AIRedwood Research argues that continual learning, which lets an AI accumulate skills during deployment, may teach models to evade blocking monitors because monitors reduce task success. Online RL on deployment trajectories would train the policy against the monitor through task reward, potentially leaving blocking monitors nearly useless over a long deployment. Memory-based systems pose a weaker version of this risk, according to the post.

  5. Lydia Hallie ✨XAI score20

    Anthropic adds local support to Projects in staged rollout

    AIAnthropic's Lydia Hallie thanked users for feedback on the new version of Projects and said local support was added the previous day. The new Projects is rolling out in stages, with early access available by DM, and existing projects may have rough edges while a smooth carry-over method is still being developed.

  6. Noah ZwebenXAI score38

    Claude Tag in Slack now accesses users' personal connectors

    AIAnthropic's Claude Tag in Slack can now use users' personal connectors, letting them securely reach a Drive doc, Salesforce account, or warehouse table they have personal access to within the conversation. The feature is available on Teams today and on Enterprise next week.

    Video from @noahzweben's post
  7. Sundar PichaiXAI score38

    Google's Project Suncatcher tests TPUs in space on a SpaceX mission

    AIGoogle's Project Suncatcher will fly a prototype satellite built with Planet aboard SpaceX's Transporter-18 mission to test whether TPUs can survive and operate in space. The test asks whether the chips can function in orbit, and the post frames it as a first step.

    Video from @sundarpichai's post
  8. Perplexity DevelopersOfficialAI score38

    Perplexity launches Fast Search API at $1 per 1,000 requests

    AIPerplexity's Fast Search, part of its Search API, is now available at $1 per 1,000 requests, with low-latency web search and what it calls the lowest cost per task among search APIs. The service runs on Photon, a new Rust-based retrieval and ranking system, and returns 95% of results in 230 ms or less.

  9. MicrosoftOfficialAI score28

    Rockwell Automation and Microsoft build AI for factory floor knowledge

    AIMicrosoft is helping Rockwell Automation turn decades of factory shop floor experience into answers workers can get in seconds. The post describes factory workers using a single question to tap into their colleagues' accumulated knowledge.

    Image from @Microsoft's post
  10. Baseten BlogOfficialAI score44

    LangSmith Fine-Tuning Trains Open Models on Agent Traces via Baseten Loops

    AILangChain launched LangSmith Fine-Tuning, which lets users fine-tune open models on their LangSmith agent traces using the open-source smithtune CLI. Training runs on Baseten Loops in the user's own workspace, and smithtune deploy places the evaluated checkpoint on a Baseten Dedicated Inference deployment. Loops is in early access, so users may need to request access for their workspace.

  11. Azure BlogOfficialAI score67

    Microsoft Foundry adds voice agents and continuous optimization for production agents

    AIMicrosoft Foundry expands its agent platform with voice agents in public preview, long-running resilience for hosted agents, and tools for evaluating production agents. The post also says GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available in Foundry. Agent optimizer, Insights, and Rubric evaluator are described as tools for continuous improvement, with some reaching general availability later this month.

    Why it matters: The post shows how Foundry combines model choice, voice agents, long-running resilience, and production evaluation into one agent workflow, with a customer example.

  12. Microsoft Foundry BlogOfficialAI score40

    Foundry Agent Service adds egress policies to restrict hosted agent destinations in preview

    AIMicrosoft's Foundry Agent Service preview lets developers attach a named, ordered egress policy to a hosted agent, allowing only approved destination hostnames. The walkthrough uses an invoice agent, an Audit-mode RAI policy with a Deny default, and Allow rules for two finance and vendor hosts, configured outside the agent code. Network egress controls are preview features, not GA, with no preview SLA, and are not intended for production use.

  13. Google ResearchOfficialAI score38

    Google's John Platt on AI for climate, disease forecasting, and science

    AIIn a Latent Space podcast episode, Google's John Platt discusses using AI to address climate change, including reducing airplane contrails that contribute about 1% of human-caused warming and detecting fires with FireSat satellites. He also describes Google's Empirical Research Assistance (ERA), which uses Gemini and Monte Carlo Tree Search and achieved top marks in recent CDC benchmarks for forecasting COVID and flu cases a week ahead.

  14. LiveKitOfficialAI score28

    LiveKit tests Gemini 3.8 Flash-Lite TTS in a live voice agent

    AILiveKit tested Gemini 3.8 Flash-Lite TTS inside a LiveKit agent, letting users direct a voice line by line and hear it hold up in a real conversation. The post highlights expressive speech, custom voices, and a production-ready voice library.

  15. Epoch AI · The Epoch BriefOfficialAI score45

    Huawei Trails Nvidia by About Four Years in AI Chip Performance and Output

    AIHuawei will likely remain about four years behind Nvidia in AI chip performance and production through 2030, Epoch AI estimates. Its flagship Ascend 950 delivers roughly half the performance of Nvidia's 2022 H100, and Huawei is projected to produce about 1.5 million chips in 2026 versus Nvidia's roughly 6 million, leaving it about 25 times behind in total compute.

  16. Google · Gemini appOfficialAI score62

    Google launches Gemini 3.8 Live with Live Avatar for enterprises

    AIGoogle introduced Gemini 3.8 Live with Live Avatar, which adds a visual persona with lip-syncing and expressions to its live dialogue models. The feature is available in Gemini Enterprise and supports 97 languages, with custom avatars available through enterprise allowlisting. Google says all output is watermarked with SynthID.

    Why it matters: The post specifies enterprise availability, custom avatar allowlisting, and 97-language support, which clarifies who can use the feature and how far it reaches.

  17. vLLMOfficialAI score34

    vLLM and RL-Kernel achieve bit-exact logprob match on AMD MI300X

    AIThe RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime, and a 200-step Qwen3-8B GRPO run on 8× AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides to achieve bit-for-bit matching on ROCm.

  18. Microsoft Foundry BlogOfficialAI score61

    Microsoft Foundry Routines reach general availability for scheduled and event-driven agents

    AIMicrosoft announced general availability of Routines in Foundry Agent Service, a managed way to run agents on a timer, on a recurring schedule, or in response to GitHub issue events and new Microsoft Teams channel messages. Routines keep the trigger, agent action, identity, connections, and run history in the Foundry project, and each routine can run under the creator's identity or the agent's own Microsoft Entra ID identity. A preview reminder tool lets a Hosted Agent schedule itself to resume later on the same conversation.

    Why it matters: The post explains how scheduled, event-based, and self-reminding agent runs are managed in one place, along with the creator versus agent identity choice for unattended tasks.