Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Aug 24

Aug 24Mon
  1. Meituan LongCatOfficialAI score23

    LongCat-2.0 now available in opencode Go for developers

    AIMeituan LongCat has made LongCat-2.0 available in opencode Go, according to the post. The post describes LongCat-2.0 as a 1.6T-parameter model with 48B active parameters, a 1M-token context window, and fully open-source release. Meituan LongCat invites users to try the model in opencode and share what they build.

  2. Qwen · new models on Hugging FaceOfficialAI score75

    Qwen3.8-Flash-Next releases open weights for a hybrid-attention architecture

    AIQwen released open weights for Qwen3.8-Flash-Next, a 125B-parameter model with 6B activated, built on a new hybrid architecture with Gated DeltaNet and Qwen Sparse Attention. The model has a native 262,144-token context length, extensible to 1,000,000 tokens, and the source reports benchmark results across coding, agent, and vision tasks.

    Why it matters: The release pairs a new hybrid attention and gated residual architecture with open weights and benchmark results, giving architecture-focused readers a concrete case to compare against prior long-context designs.

Aug 22

Aug 22Sat
  1. Ahead of AI (Sebastian Raschka)BlogAI score43

    How Claude Watermarks AI-Generated Text Through Invisible Token Sampling

    AIAnthropic plans to watermark text output from its Claude models, with the watermark invisible to users and decodable only by Anthropic. The source explains the sampling-based mechanism through a video lecture and transcript, which covers how watermarking is applied during generation and how it can fail or be removed.

Aug 21

Aug 21Fri
  1. Gemini NotebookOfficialAI score34

    Gemini upgrades Notebook for all users and adds AI Mode access

    AIGoogle's Gemini Notebook upgrade is now available to all users, with mobile support coming soon. Notebooks can also be accessed in AI Mode in Google Search, and Chat has improved math equation copy-paste and rendering, plus fixed numbers displaying backwards in right-to-left languages.

    Video from @Gemini_Notebook's post
  2. Matei ZahariaXAI score38

    Sky Lab's FreeToken runs large LLMs on consumer GPUs locally

    AISky Lab's FreeToken runs official checkpoints of large models such as Qwen3.6 35B on an 8GB RTX 4060 laptop at 39 tokens per second. The same approach reportedly serves DeepSeek-V4-Flash 284B at 22-25 tokens per second on an RTX 5090 desktop, and GLM-5.2 753B at 15 tokens per second on an RTX PRO 6000 workstation.

  3. Andrew NgXAI score31

    Andrew Ng outlines six core skills for building and deploying AI applications

    AIAndrew Ng's AI Engineering Skills Map ranks building and deploying AI applications as the top skill tier, spanning LLM foundations, data grounding, agentic systems, evaluation-driven development, production operations, and machine learning foundations. He explains that AI outputs are less predictable than traditional software, so skilled engineers build iteratively, examining results and deciding next steps based on intermediate outcomes. The skills map was derived from job postings, expert interviews, and survey responses.

  4. Jim FanXAI score59

    NVIDIA and Berkeley open-source T-Rex, a tactile robot learning method

    AINVIDIA and Berkeley are open-sourcing T-Rex, a methodology for adding touch sensing to robot manipulation models. It uses a mixture-of-transformer with a slow visuomotor expert and a fast tactile expert running four touch ticks per vision tick. A 50-hour dataset of about 5,500 episodes from 22-degree-of-freedom tactile hands is available on Hugging Face.

    Video from @DrJimFan's post
  5. DeepSeekOfficialAI score42

    DeepSeek adds vision API support via deepseek-v4-flash-vision-exp model

    AIDeepSeek's API now accepts multimodal input through the model deepseek-v4-flash-vision-exp, supporting mixed text and image requests. Each image is billed at up to 384 tokens at V4-Flash pricing, and it works with Chat Completions, Messages, and Responses endpoints. Images can be supplied as base64, external URLs, or via the Files API.

Aug 20

Aug 20Thu
  1. Ali GhodsiXAI score20

    Disaggregated storage took off after full bisection bandwidth networks emerged

    AIAli Ghodsi says disaggregating storage from compute became feasible only after research on full bisection bandwidth networks removed datacenter bottlenecks around 2010. Databricks and Snowflake followed soon after, and many others came later. He says putting data on an object store is now the standard approach.

  2. Mistral AIOfficialAI score59

    Mistral Agentic Search adds multi-step retrieval for complex enterprise documents

    AIMistral has released Agentic Search, a multi-step retrieval layer available through its Search Toolkit and Libraries. On FinanceBench, the company reports accuracy rising from 26.7% to 86% over one-shot RAG, and on OfficeQA Pro a gain from 6.3% to 51.9%. The system also reports up to 39.6% lower p90 latency and up to one-third lower token use from fewer repeated searches.

Aug 19

Aug 19Wed
  1. Ali GhodsiXAI score33

    Databricks launches AI Extract for accurate PDF field extraction

    AIDatabricks has launched AI Extract, a capability for extracting fields from PDFs that it says reaches 95% accuracy versus 87% for other tools, at very low cost. The post notes that LLMs' next-token training makes them "autocorrect" content they should preserve, which this approach is designed to avoid. The function can be called directly from SQL and used across the Databricks platform.

    Image from @alighodsi's post
  2. Jazzyear · InsightsNewsAI score29

    Jazzyear's 2026 tech investment conference maps where capital is flowing in AI and hard tech

    AIAt the 2026 Jiazi Gravity Tech Industry Investment Conference in Beijing, Jiazi Guangnian's CEO Zhang Yijia said first-half 2026 saw investment amounts rise 91.6% year on year, IPOs rise 39.2%, and M&A transaction value double. The report said AI absorbed over 70% of global venture investment, with OpenAI and Anthropic together raising $217 billion, roughly 40%.

  3. Liquid AI BlogOfficialAI score60

    Liquid AI releases DSpark draft models for LFM2.5, up to 3.2x faster inference

    AILiquid AI released DSpark speculative decoding draft models for LFM2.5-1.2B-Instruct, LFM2.5-2.6B, and LFM2.5-8B-A1B on Hugging Face. The draft models reach up to 3.18x throughput improvement on an H100 GPU and up to 2.87x on-device, and the outputs match baseline greedy decoding by construction. Support is available in llama.cpp and SGLang, with the speedup varying by model and dataset.

    Why it matters: The release reports measured speedups on both H100 and MacBook hardware, with per-dataset results and acceptance rates that show where speculative decoding helps most.

  4. Matei ZahariaXAI score46

    Databricks' custom AI Extract model reaches new frontier in document processing

    AIDatabricks says its in-house AI Extract model, paired with a custom agent harness, achieves a new frontier on complex document processing tasks. The system handles documents over 500 pages and more than 1M tokens, plus nested schemas with 1k+ objects. It decomposes large jobs, runs smaller tasks in parallel, and reconciles them into one structured output.

  5. JetBrains AI BlogOfficialAI score31

    Air Adds Multiproject View, Markdown Rendering, and Windows IME Fixes

    AIAir's latest release lets users open several projects in one window and run agents across them in parallel, with tasks grouped by project in the sidebar. Markdown files now render as formatted text while editing, with syntax shown only when editing, and Chinese, Japanese, Korean, and other IMEs now work on Windows. The release also adds a Customize screen for keymap, theme, and accent color, and lets users choose the agent and model for Agent Review.

  6. Daniel HanXAI score40

    Unsloth releases 1-bit Qwen3.8-27B quants running on 8GB RAM

    AIUnsloth has released 1-bit quantized versions of Qwen3.8-27B that run on 8GB of RAM while retaining about 77% of BF16 accuracy. The team originally hesitated to publish them but was surprised by how well they performed in internal testing. The release accompanies new Qwen3.8-27B GGUFs that the company says deliver 10% higher accuracy.

  7. Daniel HanXAI score40

    Unsloth releases Qwen3.8-27B GGUFs with Dynamic v3 quantization

    AIUnsloth released new Qwen3.8-27B GGUF quantizations built with Unsloth Dynamic v3, which it says gain about 10% top-1% accuracy at the same size. The accuracy was measured with the new Divergence-300 metric, which extends top-1% greedy accuracy to 32 tokens using 300 unseen examples from Terminal Bench and DeepSWE. Unsloth also released 1-bit quants that it says run in 6–8GB, with 8GB RAM cited for running them.

  8. Google · new models on Hugging FaceOfficialAI score22

    TIPS So400m/14 v1 Vision-Language Model Released on Hugging Face

    AIGoogle released google/tipsv1-so400m14, the original v1 So400m/14 checkpoint of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 413M vision parameters and 448M text parameters at 448 resolution, and is licensed under Apache 2.0.

Aug 18

Aug 18Tue
  1. Liquid AI BlogOfficialAI score65

    Liquid AI releases QAD 4-bit LFM2.5 checkpoints for edge deployment

    AILiquid AI released 4-bit Q4_0 GGUF checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B, trained with Quantization-Aware Distillation. The company says the checkpoints recover most accuracy lost to quantization, reaching roughly 97% of their BF16 averages while keeping Q4_0 memory footprint and throughput. Benchmarks compare them against post-training quantized Q4_0 GGUFs and against Q5_K_M, Q4_K_M, and Unsloth's UD-Q4_K_XL.

    Why it matters: The post shows how quantization-aware distillation recovers accuracy lost in Q4_0 checkpoints, with throughput measured across four hardware backends for deployment tradeoffs.

  2. Jazzyear · InsightsNewsAI score47

    Unitree Lists on STAR Market, Opens at 1,100 Yuan, Up 629% From IPO Price

    AIUnitree Robotics debuted on the Shanghai Stock Exchange STAR Market at 1,100 yuan per share, 629.44% above its 150.8-yuan issue price, with a total market value of 444.9 billion yuan. The company's founder, Wang Xingxing, has been known for taking unconventional positions, including low-cost in-house core components and a skeptical view that data alone will not produce embodied intelligence.

  3. Cursor ChangelogOfficialAI score62

    Cursor adds event subscriptions, custom modes, and subagent VMs for cloud agents

    AICursor's update lets cloud agents subscribe to PRs, Slack threads, and scheduled tasks, and wake when something happens. It also adds custom modes that pin a skill in chat, subagents that run on their own virtual machines, and a /goal command for long-lived objectives. Users can also send steering messages while an agent works, with follow-ups applied at the next tool call.

    Why it matters: The release lists concrete agent controls such as event subscriptions, custom modes, subagent VMs, and /goal, showing how cloud agents may run longer tasks with less manual steering.

  4. OpenRouter BlogOfficialAI score72

    OpenRouter announces it is joining Stripe, keeping its product unchanged

    AIOpenRouter announced it is joining forces with Stripe, saying its product, name, mission, and roadmap will remain the same. The company says it processes more than 10 trillion tokens per day from over 400 AI models for a community of over 10 million developers and companies. The transaction is subject to customary closing conditions and is expected to close in the coming weeks.

    Why it matters: The announcement states that OpenRouter's product, roadmap, and mission stay unchanged after the Stripe deal, which clarifies what existing developers should expect.