Midjourney unveils technical dive into its new Scanner
AIA technical dive inside our new "Midjourney Scanner"
Updated
Updated
AIA technical dive inside our new "Midjourney Scanner"
AIMidjourney has released a big-batch draft mode for V8.1 that generates 24 lower-resolution images at half the price of a standard-resolution four-image job. Users can press "Vary" on any image they like to get new full-resolution versions.
AIByteDance has released Sa2VA-LLaVA-1.5-7B on Hugging Face, a model built on LLaVA-1.5-7B with a SAM2 grounding encoder that performs dense image and video referring segmentation alongside open-ended chat. The checkpoint is self-contained and loads with trust_remote_code=True without extra packages, and it is positioned as a LISA-comparable baseline within the Sa2VA family. Reported results include 80.3 cIoU on RefCOCO val and 54.8 J&F on MeViS (val_u).
AIByteDance's Sa2VA-Qwen3-VL-4B-SAM3 is built on Qwen3-VL-4B-Instruct with a SAM3 grounding encoder and produces dense image and video referring segmentation alongside chat. It reports 83.7 cIoU on RefCOCO val, 65.3 J&F on MeViS (val_u), and 77.1 on Ref-DAVIS17. The checkpoint is self-contained and loads on Hugging Face with trust_remote_code=True, with no extra packages required.
AIStability AI says Stable Audio 3.0 was built to support exploratory audio work such as stem remixing. A quoted post from @teropa reports using the Medium model with an init_audio input and an init_noise_level of 0.4–0.5, with empty prompts.
AIFei-Fei Li praised World Labs for partnering with Lore to turn creative ideas into interactive experiences for users. A linked World Labs post describes worlds running directly in a browser rather than as video, populated with history's greatest minds.
AIZ.ai released SCAIL-2, an open-source model that animates a reference character from a driving video without skeleton maps or inpainting masks. It also supports character replacement, multi-character scenes, and animal-driving, with 512p and 704p resolutions and inputs whose height and width are both divisible by 32.
AIByteDance has open-sourced the 1.3B-parameter weights of its Bernini Renderer (Bernini-R), available on Hugging Face as ByteDance/Bernini-R-1.3B-Diffusers. Fine-tuned from Wan2.1-1.3B, the model performs close to the 14B variant on simple tasks such as style transfer, subtitle or watermark removal, and local editing, but lags on complex tasks such as human generation. The release requires a CUDA GPU, with an H100 recommended for FlashAttention-3.
AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.
AIPaddleOCR and ERNIE Image are now available as official Dify plugins, bringing document parsing and image generation into Dify's agent workflows. PaddleOCR, powered by PP-OCRv5, PP-StructureV3, and PaddleOCR-VL, turns images, scanned PDFs, and multilingual documents into structured data for chunking, vectorization, and RAG, with private or on-prem deployment supported. ERNIE Image offers free generation, a Turbo mode with 8-step inference, and an OpenAI-style API.
AIFei-Fei Li praised GPIC, a new benchmark dataset for visual generation built for modern large-scale generative models. The corpus includes 100M VLM-captioned image-text pairs for training and 1M pairs for benchmarking, totaling about 28 trillion pixels. It is centrally hosted and fully permissive for research and commercial use.
AIStability AI has launched Stable Audio 3.0 for experimentation. The Small and Medium models are available on Hugging Face, while the Large model is accessible through the Stability AI API or self-hosting with an enterprise license.
AIOpenAI's Nick Turley says ChatGPT images are useful across many professional settings, not just as a novelty. Ethan Mollick, citing weeks of use with GPT ImageGen-2, reports the quality now reliably produces readable text, slides, and academic-paper-style figures.
AINick Turley of OpenAI reacted with a one-word "good!" to Arena's report that GPT-Image-2 ranked first across all Image Arena leaderboards. Arena cited a 1512 Text-to-Image score, a 242-point lead over the second-place Nano-banana-2 model.
AINick Turley announced ChatGPT Images 2.0 as a major advance in image generation, citing better adherence to detailed instructions, rendering of dense text, and more accurate understanding of the world. He said the model can spend extra time planning and refining outputs for tasks needing more accuracy and clarity, and that users have generated over 1 billion images with ChatGPT.
Why it matters: The post names concrete gains in instruction following, dense text rendering, and optional extended thinking for image output, which helps readers gauge practical scope.
AIResearchers introduced a simple, sample-efficient online reinforcement learning technique for post-training image generation models. It is presented as a possible steerable alternative to classifier-free guidance (CFG) that can be driven by any scalar reward, including human preference.
AIStability AI has introduced Brand Studio, an end-to-end creative production platform for enterprise teams that builds around each brand's identity. Its Brand Central hub supports custom Brand ID models and Campaigns, while Producer Mode turns prompts into step-by-step production plans. Curated Model Routing selects models including Stable Diffusion, Nano Banana, and Seedream, and new Precision Inpainting and Product Insertion tools enable targeted edits.
AIBlack Forest Labs released FLUX.2 Small Decoder, a distilled VAE decoder that works as a drop-in replacement for the standard FLUX.2 decoder on Hugging Face. It decodes about 1.4x faster and uses about 1.4x less VRAM at decode time, with ~28M decoder parameters versus ~50M in the full decoder and minimal quality loss. It is available under the Apache 2.0 license and is compatible with FLUX.2-klein-4B, FLUX.2-klein-9B, FLUX.2-klein-9b-kv, and FLUX.2-dev.
AIInternVL-U is a lightweight 4B unified multimodal model that combines reasoning, generation, and editing in one framework, according to Intern Large Models. The post says it uses unified contextual modeling, modality-specific modular design, and decoupled visual representations to balance performance and efficiency. It reportedly outperforms unified baselines more than 3× its size on text rendering, scientific reasoning, and spatially grounded generation and editing, and is open-source on GitHub and Hugging Face.
AIBlack Forest Labs has released FLUX.2 [klein] 9B-KV, a variant of FLUX.2 [klein] 9B that caches reference-image key-value pairs to speed up multi-reference editing by up to 2.5 times. The 9B flow model, which uses an 8B Qwen3 text embedder and is step-distilled to 4 inference steps, is available for non-commercial use under the FLUX Non-Commercial License and fits in about 29GB VRAM.
AIMeta is working with environmental survey teams to use its Segment Anything Model (SAM) to identify minute changes in water conditions in satellite images. The company hopes faster analysis will support rapid flood response and keep communities safer.
AINano Banana 2, officially released as Gemini 3.1 Flash Image Preview, ranks first in Image Arena text-to-image with a score of 1279. The quoted post says it also ties for first in single-image editing at 1407 and costs $0.067 per image, about half the price of Nano Banana Pro.
AIBlack Forest Labs released FLUX.2 [klein] 4B, a 4 billion parameter model that unifies text-to-image generation and image editing with multi-reference support. The source says it runs on consumer GPUs such as the RTX 3090 or 4070 with about 13GB VRAM, and its open weights are available under the Apache 2.0 license.
Why it matters: The source specifies a 4 billion parameter model running on about 13GB VRAM under Apache 2.0, which helps readers judge whether local image generation fits their hardware.
AIBlack Forest Labs has published FLUX.2 [klein] 4B Base, a 4 billion parameter text-to-image model that also supports multi-reference editing. The model is undistilled, is released with open weights under Apache 2.0, and is described as fitting in about 13GB VRAM on cards such as the RTX 3090 or 4070, with reference code available in its GitHub repository and support in ComfyUI and Diffusers.
AIBlack Forest Labs has released FLUX.2 [klein] 9B Base, a 9 billion parameter undistilled rectified flow transformer with open weights for text-to-image generation and multi-reference editing. The model is intended for fine-tuning, LoRA training, and research, and fits in about 29GB VRAM on NVIDIA RTX 4090-class GPUs. A reference implementation is available on GitHub, and the model works with ComfyUI and Diffusers.
AIZ.ai has launched GLM-Image, an image generation model built on a multimodal architecture that combines autoregressive semantic understanding with diffusion-based decoding. The update improves knowledge-intensive generation and makes text rendering inside images more stable and accurate, suiting commercial design and educational illustrations.
AITsinghua AIR and BAAI's DrugCLIP screened 10,000 proteins against 500 million molecules, identifying over 2 million drug candidates. The post claims a 1-million-fold speedup, reaching 10 trillion protein-molecule pairs per day, and positions DrugCLIP as bridging AlphaFold structures to drug candidates. The work is published in Science, with a platform available at drugclip.com.
AIOpenAI's new ChatGPT Images is rolling out in ChatGPT starting today. The update offers more precise edits, stronger instruction following, and up to 4x faster generation while preserving lighting, composition, and likeness across edits.
AIApple has released SHARP, a model that generates a 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The output renders in real time as high-resolution photorealistic views of nearby camera positions, with metric absolute scale, and the paper reports reductions of 25–34% in LPIPS and 21–43% in DISTS versus the best prior model.
AIQuoc Le says Nano Banana Pro's text understanding is impressive, showing an infographic it created for the sequence-to-sequence paper from a single request. The post includes the generated image as its main evidence, with no figures on speed, pricing, or benchmarks.
AIByteDance's R&D team created the technology behind TikTok's "Cartoonify" sticker in just 48 hours during its annual Hackathon. The augmented reality effect makes physical objects on screen dance, cry, sing, and perform other actions.