Skip to contentSkip to stories

Updated

#Image generation

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. MarkTechPostNewsAI score62

    Alibaba Qwen releases Qwen-Image-2.1-Turbo, an 8-step 7B image model

    AIAlibaba's Qwen team released Qwen-Image-2.1-Turbo, an accelerated checkpoint of Qwen-Image-2.1 that generates and edits images in 8 denoising steps instead of 40. The model keeps the same 7B architecture and offers a hosted API at CNY 0.1 per image, while its weights are under a Qwen Research License that requires separate permission for commercial self-hosting.

  2. Qwen · new models on Hugging FaceOfficialAI score49

    Qwen releases Qwen-Image-2.1-Turbo, an 8-step accelerated image generation checkpoint

    AIQwen has published Qwen-Image-2.1-Turbo on Hugging Face, an accelerated checkpoint of Qwen-Image-2.1 for text-to-image generation and image editing with 8 denoising steps. The checkpoint uses the same 7B visual generation architecture, loads directly with QwenImage21Pipeline in Diffusers, and includes its recommended sampling schedule. It defaults to CFG=1 and uses prefix KV caching to reuse text and reference-image context across steps.

  3. ModelScopeOfficialAI score60

    Qwen-Image-2.1-Turbo cuts image generation and editing to 8 denoising steps

    AIModelScope announces Qwen-Image-2.1-Turbo, an accelerated checkpoint that keeps the 7B visual architecture and runs image generation and editing in 8 denoising steps. The source says it uses CFG=1 and prefix KV caching to reuse text and reference-image context across steps, supports 2048 resolution with square, portrait, landscape, and widescreen presets, and loads through QwenImage21Pipeline in Diffusers. It is released under the Qwen Research License Agreement.

    Why it matters: The source names a concrete speedup path, 8 sampling steps and CFG=1 with prefix KV caching, which matters to anyone weighing image generation latency.

    Image from @ModelScope2022's post
  4. LeiphoneNewsAI score42

    Doubao Work adds Canvas feature and Doubao 2.1 Lite model

    AIDoubao Work has added a Canvas feature for complex creative tasks, placing materials, design plans and outputs on one infinite canvas where users can keep editing text, colors and layout after images are generated. The update also integrates the lightweight Doubao 2.1 Lite model, aimed at everyday Q&A, document writing, spreadsheets and PPT creation, with optimized response speed and usage consumption.

Oct 8

Oct 8Thu
  1. LumaOfficialAI score22

    Jon Erwin's Moses made using AI to extend real actors' worlds

    AILuma Labs posted that filmmaker Jon Erwin made Moses on a Manhattan Beach stage with real actors, using AI to carry them into any world the story required. The post frames AI as removing budget limits on where a story can be filmed.

    Video from @LumaLabsAI's post
  2. Artificial AnalysisOfficialAI score31

    Grok Imagine Video 1.5 Lite leads on quality and speed benchmark

    AIAmong 12 models on AA-Video-T2V-Silent v2.0, Grok Imagine Video 1.5 Lite is the only one that is both fastest and highest quality, with no model beating it on both measures. It generates a 10-second 1080p clip in a median of 60.5 seconds. Kling 3.0 1080p (Pro) scores slightly higher but takes 94 seconds for a 5-second clip, while Vidu Q3 Turbo is 9 seconds faster on a 5-second 720p clip yet scores well below it.

    Image from @ArtificialAnlys's post
  3. MidjourneyOfficialAI score34

    Midjourney tests a "thinking mode" for image generation on Alpha

    AIMidjourney is testing a new "thinking mode" for its image generation on its Alpha website, alpha.midjourney.com. The company says the mode improves prompt accuracy, typography, and coherence, and it is asking users to try it and share feedback.

    Image from @midjourney's post
  4. Midjourney UpdatesOfficialAI score46

    Midjourney Tests Thinking Mode for Image Generation on Alpha Site

    AIMidjourney is testing a "Thinking Mode" on its Alpha website, where users can click "Rerun (Thinking)" in a job's lightbox to regenerate an image. The company says early tests show gains in prompt accuracy, typography, and coherence, and it is asking users to share feedback in its #ideas-and-features channel. It may later offer the mode broadly or as an option to add more thinking after a job.

  5. Midjourney UpdatesOfficialAI score25

    Midjourney Adds Shared Folders and Thinking Mode in Alpha Update

    AIMidjourney's alpha site now lets users share folders with others as Collaborators or Viewers, with sharing by link also available. A new thinking mode lets users rerun jobs made with 8.2 standard and edit models to fix missed prompt details such as objects, layout, anatomy, and text. Sharing does not change image privacy, so non-stealth images can still appear on Explore and profiles.

  6. RunwayOfficialAI score36

    Runway promotes Claude Motion for animating charts and explainers

    AIRunway's post promotes using Claude Motion to animate charts, create customer walkthroughs, and make short explainers. The animated work can then be brought into Runway to generate videos and images. Claude Motion is described as being in beta, per the quoted Claude post.

    Video from @runwayml's post
  7. OpenRouterOfficialAI score40

    Grok Imagine Video 1.5 Lite now available on OpenRouter

    AIOpenRouter now offers xAI's Grok Imagine Video 1.5 Lite for text-to-video and image-to-video generation. The quoted post from Grok Imagine lists pricing of $0.02 per second at 480p, $0.03 per second at 720p, and $0.14 per second at 1080p.

  8. Vercel DevelopersOfficialAI score26

    FLUX 3 Image is now available on Vercel AI Gateway

    AIVercel says FLUX 3 Image, from Black Forest Labs, is live on AI Gateway for generating and editing images. It supports up to 10 reference images and 4K output.

    Image from @vercel_dev's post
  9. ClaudeOfficialAI score46

    Claude Motion turns reports into editable code-based animations

    AIAnthropic's Claude Motion converts reports, charts, or product walkthroughs into short animations. Claude writes each animation as code rather than using a video model, so users can edit any word, number, or timing and export an MP4. The feature is in beta on Team and Enterprise plans.

    Image from @claudeai's post
  10. QbitAINewsAI score47

    Vidu Q4 Preview Offers 4K Video Generation at About 0.09 Yuan per Second

    AIShengshu Technology has opened a preview of its Vidu Q4 video generation model, which supports native 4K output and up to 15 reference images and three reference audio clips. Testers generated a one-minute video for about 5.4 yuan, roughly 0.09 yuan per second at 720P, which the article says is a starting price that varies by resolution and mode. The Vidu Q4 preview is available through the Vidu platform, with the MaaS API priced at about 0.6 yuan per second for 720P image-to-video.

Oct 7

Oct 7Wed
  1. Apple Machine Learning ResearchOfficialAI score42

    Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

    AIApple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.

  2. Amazon ScienceOfficialAI score35

    Amazon's AI smart glasses guide delivery drivers hands-free to doorsteps

    AIAmazon has developed AI-powered smart glasses that guide delivery drivers from their van to the customer's doorstep without using their hands. Amazon applied scientist Yelin Kim will discuss the computer vision and edge AI behind the system at COLM 2026.

    Video from @AmazonScience's post
  3. 🚨 AI News | TestingCatalogXAI score34

    Envato launches Burst Mode, generating up to six visual directions per credit

    AIEnvato has launched Burst Mode, which turns a single idea or reference image into up to six visual directions for one AI credit, up to 10x faster. Users can steer each batch with references, moodboards, and a creativity control ranging from focused to wild, then refine promising results with "More like this."

    Video from @testingcatalog's post
  4. elvisXAI score22

    Envato launches Burst Mode for AI image direction exploration

    AIEnvato has launched Burst Mode, which starts from a rough idea or reference image and generates up to six image directions side by side. Users can steer those directions with moodboards and a creativity control ranging from focused to wild. Envato says it can produce up to six directions at once, up to 10× faster, for one AI credit.

  5. falOfficialAI score34

    Vidu Q4 video model now live on fal with native audio

    AIfal has launched Vidu Q4, offering image-to-video generation from a single first frame with native audio. Reference-to-video supports up to 12 reference images and 3 voice clips for consistent characters and voices. Clips run 3 to 16 seconds at resolutions from 540p up to 4K.

    Video from @fal's post
  6. SantiagoXAI score32

    Envato's Burst mode returns six image directions from one idea

    AIEnvato's new Burst mode turns a single idea into six image directions at once, for one AI credit, which the company says is up to 10× faster. Users can then refine a direction with "more like this" or pick a specific style with "polish."

  7. Google FlowOfficialAI score22

    Nano Banana 2.1 renders subtle text within a monstera leaf photo

    AIGoogle's Nano Banana 2.1 image model rendered the words "Nano Banana" organically within the natural gaps of a monstera leaf in a photo, as shown by @fofrAI. The post highlights the model handling a subtle prompt involving nuanced text placement.

  8. Google FlowOfficialAI score22

    Nano Banana 2.1 shows improved prompt comprehension in Google Flow tests

    AIGoogle Flow's Nano Banana 2.1 is reported to have improved prompt comprehension, with a first test of the prompt "3 centaurs doing backflips" producing a shot the tester called perfect. The claim rests on a single informal example from the tester, @HarshithLucky3, and no benchmark or broader evaluation is cited.

  9. Google FlowOfficialAI score27

    Google's Nano Banana 2.1 image model is already being tested

    AIGoogle says users are already testing its newest image generation model, Nano Banana 2.1, and shares examples of the early results. The post provides no benchmark scores, pricing, or capability details.

  10. Google FlowOfficialAI score23

    Google Flow highlights strong consistency across repeated edits in new model

    AIGoogle Flow says its model excels at maintaining consistency after multiple edits. A test by @pleometric, reconstructing an image 100 times with the prompt "reconstruct this exactly as it is," found 2.1 best preserved the original subject compared with other Nano Banana models.

  11. Google DeepMind · The KeywordOfficialAI score62

    Google expands SynthID Detector globally to check AI-generated media

    AIGoogle is making its SynthID Detector available globally in English, letting anyone check whether an image, video, or audio file was made with AI from Google or partners including OpenAI, NVIDIA, Kakao, and soon Apple. The tool joins built-in verification in Search, the Gemini app, and Chrome, which now handle over 1 million requests daily. Google says SynthID has watermarked over 180 billion images and videos and 240,000 years of audio.

    Why it matters: The source specifies which vendors' AI media the detector checks, helping readers judge how far the verification covers content they encounter online.

  12. PixVerseOfficialAI score20

    PixVerse plugin lets users generate videos directly within chat

    AIThe PixVerse plugin enables video creation from text, images, or video references inside the chat interface. Users select the plugin, describe a scene or add a reference, then specify model, duration, resolution, and aspect ratio before generating.

    Video from @PixVerse's post
  13. laurenXAI score31

    Grok Bot to route tasks to best third-party models

    AIGrok Bot will now use the best backend model for each task, including Claude Opus 5.5, MidJourney, Suno, and other leading APIs. The change is framed as choosing whatever is most likely to produce the best outcome for users.

Oct 6

Oct 6Tue
  1. Josh WoodwardOfficialAI score34

    Nano Banana 2.1 adds mask-based editing and improved visual quality

    AIGoogle's Nano Banana 2.1 is an upgraded image model that outperforms prior versions in visual design, mask-based editing, subject consistency, and natural-looking imagery. Josh Woodward calls mask-based editing his favorite feature from the launch and says more is coming soon.

  2. meng shaoXAI score52

    xAI Cookbook adds five apps, expanding Grok API examples to ten

    AIThe xAI Cookbook now has ten runnable Grok API examples across three tracks: real-time voice agents, multimodal generation, and live X data analysis. The author says four voice examples show the same Realtime Voice API across WebSocket, WebRTC, Twilio phone, and mobile transports. The four multimodal examples chain understanding, image generation or editing, video, and TTS, with Grok making creative decisions and Imagine models executing them.

    Image from @shao__meng's post
  3. ComfyUIOfficialAI score21

    ComfyUI announces Gemini Nano Banana 2.1 availability

    AIComfyUI says Gemini Nano Banana 2.1 is now available, linking to a blog post with details. The post itself provides no further specifics about features, pricing, or capabilities.

  4. ComfyUIOfficialAI score34

    Nano Banana 2.1 arrives in ComfyUI via Partner Nodes

    AIComfyUI announces that Nano Banana 2.1 is now available through Partner Nodes. The model supports 1K to 4K output, Minimal, Medium, and High thinking levels, and up to 14 reference images. It also renders text exactly as written and supports targeted, multi-turn edits.

    Video from @ComfyUI's post
  5. Comfy BlogOfficialAI score43

    Gemini Nano Banana 2.1 is now available through ComfyUI Partner Nodes

    AIGoogle's Gemini Nano Banana 2.1 image generation and editing model is now available in ComfyUI through Partner Nodes, the successor to Nano Banana 2. The model accepts a prompt and up to 14 reference images, outputs at up to 4K, and offers Minimal, Medium and High thinking levels. It adds a 9:21 aspect ratio and, according to the post, costs less per run than Nano Banana 2.

  6. falOfficialAI score34

    fal Now Available in ChatGPT and Codex for Media Generation

    AIfal is now available inside ChatGPT and Codex, letting users generate images and videos without leaving the chat. Generated media can be browsed directly in the conversation, and users can access their fal Media Library from ChatGPT. The library can be pinned to the sidebar for quick access to assets in any chat.

    Video from @fal's post