Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Sep 28

Sep 28Mon
  1. Unsloth AIOfficialAI score34

    Laya Decision models can now run locally on 4GB RAM

    AIUnsloth AI says Laya Decision models can run locally on just 4GB of RAM, on CPU, Mac, Windows, Linux, and GPU setups. The post adds that Laya can be served through a Jev-compatible API via Unsloth Desktop.

    Image from @UnslothAI's post
  2. ModelScopeOfficialAI score46

    Qwen-Image-2.1 LoRAs extract and remove layers for editing

    AIModelScope released two Qwen-Image-2.1 LoRAs, LayerExtract and LayerRemove, for layer-based image editing. LayerExtract isolates a prompt-specified subject onto a transparent background, while LayerRemove deletes the matching object from the source image and reconstructs the scene behind it. Both can be hot-swapped within the same DiffSynth-Studio pipeline, and the LoRA weights are licensed under Apache 2.0, with Qwen-Image-2.1 base-model terms also applying.

    Image from @ModelScope2022's post
  3. ModelScopeOfficialAI score43

    Jina-OCR-v1 parses full pages into Markdown at 2.57 pages per second

    AIJina-OCR-v1, a 3.4B-parameter MoE model that activates 570M parameters per token, converts entire document pages into structured Markdown at 2.57 pages per second. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter, and delivers the highest throughput among 14 evaluated systems at concurrency 32. The model is released under CC BY-NC 4.0, so commercial use requires permission.

    Image from @ModelScope2022's post
  4. Mastra BlogOfficialAI score49

    Mastra Adds Classifiers for Choice, Score, and Boolean Decisions

    AIMastra now offers classifiers that use evaluation models to answer questions defined as choice, score, or boolean, returning criteria keys, ordered positions, or true probabilities. Classifiers are registered on the Mastra instance and can drive workflow branching, such as routing a request to one of several agents. The feature requires @mastra/core 1.69.0 or later.

Sep 26

Sep 26Sat
  1. DeedyXAI score18

    Claude generated a 2-minute video on Google's history

    AIDeedy shared a two-minute video about the history of Google that was generated by Claude. The post calls the result impressive but gives no details on the tools used or the video's production.

    Video from @deedydas's post
  2. OpenCodeOfficialAI score2

    OpenCode posts "day two" update with no product details

    AIOpenCode's post is a brief teaser reading "it takes two to tango" and "day two," with no product, feature, or figure specified. The post offers no concrete details about what was announced or changed.

    Video from @opencode's post

Sep 25

Sep 25Fri
  1. KreaOfficialAI score22

    Krea launches an agent for creative workflows

    AIKrea published a post linking to presenting an agent-focused product page. The post itself provides no further details on features, models, pricing, or availability.

  2. Alex HeathXAI score18

    Warp unveils Warp 2.0, an AI agent for HR and payroll tasks

    AIWarp has launched Warp 2.0, which it calls the first AI Head of HR, backed by $85M in funding. The company positions its Warp Agent to handle administrative HR work such as state tax registration and payroll compliance, so people can focus on managers and company culture.

  3. SemiAnalysisBlogAI score59

    China Holds Over 24GW of Datacenter Capacity, Shifting Inland With AI Demand

    AISemiAnalysis's China Datacenter Model tracks over 1,000 facilities across 60+ operators and puts China's fleet above 24GW, larger than EMEA. The report attributes the buildout to the Eastern Data, Western Compute policy and hyperscale AI demand, which is moving capacity to western hubs such as Inner Mongolia at construction speeds it says the West cannot match.

  4. OpenCodeOfficialAI score22

    OpenCode makes $60 of DeepSeek v4.1 Flash usage permanent

    AIOpenCode says its $60 of usage for DeepSeek v4.1 Flash is now permanent, as part of its "Operation Cheepseek Phase 2" promotion. The post gives no further details on terms, duration, or eligibility.

  5. LlamaIndex 🦙OfficialAI score18

    LlamaParse preserves complex Fed forecast tables for AI analysis

    AILlamaIndex tested LlamaParse on the Fed's September 2026 projections PDF, where a 2029 column's June comparison cell is blank. The company says all nine GDP median figures on page 2 matched the original PDF after parsing, with alignment kept in the returned HTML.

    Image from @llama_index's post

Sep 24

Sep 24Thu
  1. ModelScopeOfficialAI score23

    NeoHorse-Jev-4B open model turns app states into structured decisions

    AIModelScope has released NeoHorse-Jev-4B, a compact open model that converts application states into structured decisions and probabilities. It scores 77.70 across six text decision benchmark groups, ranking first among four open-weight models with complete results in the comparison. Its prefill-only inference supports Choice, Noul, and Score primitives, accepts text or a single image with text, and is available under Apache 2.0 for deployment via vLLM, SGLang, Python, CLI, or HTTP.

    Video from @ModelScope2022's post
  2. vLLMOfficialAI score30

    vLLM integrates TileRT with PD disaggregation, benchmark config published

    AIvLLM has published a blog post explaining how its integration with TileRT works, alongside a public benchmark configuration in the InferenceX repository. The benchmark script covers GLM-5.3 FP8 on MI355X hardware and is linked on GitHub. The post itself provides the integration details.

  3. vLLMOfficialAI score42

    TileRT and vLLM hit 469 tok/s on GLM-5.3 with MI355X

    AIThe TileRT and AMD teams reached 469 tok/s single-user decode for GLM-5.3 on 8× MI355X using vLLM. The setup disaggregates work, with vLLM handling prefill and TileRT handling latency-critical decode through vLLM's V1 connector interface. SemiAnalysis's AgentX benchmark reports the configuration at 470 TPS on GLM 5.3 (FP8), over 40% faster than GB300 TRTLLM using FP4.

  4. WorkBuddyOfficialAI score10

    WorkBuddy announces a campus meetup on September 27

    AIWorkBuddy invited Penn students to a campus meetup on September 27, encouraging attendees to bring a friend. The post provides no further details about the event's format, agenda, or location.

    Image from @WorkBuddy_AI's post
  5. WorkBuddyOfficialAI score4

    WorkBuddy promotes its help with North American job hunting steps

    AIWorkBuddy promotes its assistance across North American job hunting, including resume refinement, interview preparation, and comparing job offers. The post is promotional and gives no specific features, figures, or availability details.

    Video from @WorkBuddy_AI's post
  6. vLLMOfficialAI score34

    vLLM and RL-Kernel achieve bit-exact logprob match on AMD MI300X

    AIThe RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime, and a 200-step Qwen3-8B GRPO run on 8× AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides to achieve bit-for-bit matching on ROCm.

  7. OpenCodeOfficialAI score60

    OpenCode Server has a code execution vulnerability in versions 1.14.30 through 1.18.21

    AIOpenCode warned that a code execution vulnerability affects OpenCode Server versions 1.14.30 through 1.18.21 and urged users to update to the latest version. The post credits @christophetd and the Datadog team for reporting the issue, with details linked in a Datadog Security Labs article.

    Why it matters: The post names the affected version range and the fix, which matters for anyone running OpenCode Server and deciding whether to upgrade now.

  8. OpenBMBOfficialAI score34

    FIT-GGUF enables size-targeted mixed-precision quantization of MiniCPM5-2B

    AIDeveloper @Scorp1o_117 used FIT-GGUF to build four MiniCPM5-2B GGUF variants, ranging from about 1.14 GiB to 1.46 GiB, tuned to target file sizes or fidelity tiers. Instead of fixed presets, FIT-GGUF allocates precision tensor by tensor, with Quality, Balanced, Compact, and Mini options, and its generated files matched predicted sizes. Builds are evaluated with KL Divergence and Same-top metrics and are available on Hugging Face.

    Image from @OpenBMB's post
  9. KrASIA · Big TechNewsAI score55

    Mind Lab launches Mint Recursive, a post-training platform for companies

    AIMind Lab unveiled Mint Recursive, a post-training and inference platform for industry use, alongside Macaron-V1.1, a model post-trained entirely on it. Macaron-V1.1 is a 752-billion-parameter model built from GLM-5.3 with four two-billion-parameter LoRA expert modules for chat, agents, coding, and generation. The platform is serverless and bills by token usage, and it collects feedback from models in use to support continued training.

Sep 23

Sep 23Wed
  1. OpenClaw🦞OfficialAI score10

    OpenClaw plugins can use optional models for structured decisions

    AISupporting plugins can call an optional model for structured choices, kept separate from chat. TypeSafe Jev sends supplied information to its hosted API and incurs normal charges, while ONNX offers local CPU options. Decision Models are off by default.

  2. OpenClaw🦞OfficialAI score12

    OpenClaw adds live meeting notes and saved transcript tabs

    AIOpenClaw lets users follow notes while a Google Meet, Teams, Zoom, or voice capture is still running. The saved transcript can be opened in its own tab, though transcription may lag. Generated notes use the user's configured model and incur its usual charges.

  3. OpenClaw🦞OfficialAI score16

    OpenClaw Usage view shows 30-day totals and who started sessions

    AIOpenClaw's Usage view opens on the last 30 days and includes totals beyond the rows shown on screen. Its Started by view shows who began the sessions behind tokens, estimated costs, and counts. The post notes that costs are estimates, not bills.

  4. Amjad MasadXAI score5

    Amjad Masad posts an eyes emoji teaser with no details

    AIReplit CEO Amjad Masad posted only an eyes emoji, with no product name, feature, or figure stated. The post is most likely a teaser reacting to a related post by @MakerThrive, which promises a launch in less than a week built entirely on Replit.

  5. vLLM BlogOfficialAI score54

    vLLM adds distortion-free Gumbel-max watermarking for text provenance

    AIvLLM now supports Gumbel-max watermarking, which embeds a keyed signal into generated text without changing the expected token distribution. Detection requires the secret key and tokenizer, and the signal accumulates over longer outputs. Benchmarks on Qwen3.5-27B with MTP-3 show throughput changes between -1.1% and +2.0% across batch sizes, with no consistent slowdown.

  6. SemiAnalysisBlogAI score85

    SemiAnalysis releases ClusterMAX 3.0, rating 77 GPU clouds through hands-on testing

    AISemiAnalysis releases ClusterMAX 3.0, a rating of managed GPU clusters from neoclouds that covers 77 providers, with 323 in its market view. The rating is based on audit, performance, and reliability tests, along with interviews with over 200 end users. CoreWeave and Nebius hold the Platinum tier, Google Cloud and Oracle hold Gold, and only 19 providers earned a Medallion rating.

    Why it matters: The report shows how GPU cloud providers are ranked through hands-on tests, with details on benchmarks, reliability checks and SLA terms that buyers can reuse.

  7. InferactOfficialAI score49

    Inferact's TPU megakernel runs Kimi K3 at 709 tokens/s

    AIInferact says its first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode with DSpark speculative decoding, versus 450 tokens/s for its GB200 baseline. The company claims it is the first TPU inference megakernel, running the whole model in a single Pallas kernel, and says it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8 without speculative decoding. Inferact says it is open-sourcing the kernel today.

    Video from @inferact's post
  8. LM StudioOfficialAI score28

    Bionic adds a built-in interactive canvas for shared diagrams

    AIBionic now includes a built-in interactive canvas where users can create Excalidraw diagrams that both they and Bionic can view and edit. The canvas supports collaboration on mockups, system designs, and process maps, and users can ask Bionic to implement what is drawn.

    Video from @lmstudio's post