MultiModal COnnection with Representation Aligned Latent Embeddings — a unified framework for multimodal image generation and editing that bridges VLM reasoning with diffusion synthesis.
About the Project
MMCORE leverages a pre-trained Vision-Language Model (VLM) to predict semantic visual embeddings via learnable query tokens, which serve as conditioning signals for a diffusion model. This streamlined design effectively transfers the rich understanding and reasoning capabilities of VLMs into the visual generation process, reducing computational overhead while maintaining high-fidelity synthesis.
Visual query tokens capture global semantics while text embeddings preserve fine-grained lexical details.
Distillation from frozen ViT encoders provides stable supervision for learning robust visual tokens.
Minimal overhead through efficient alignment of pretrained models, avoiding costly end-to-end retraining.
Scales to 10+ input images with fine-grained localized control for complex multi-image editing.
Architecture
Multimodal information from a VLM is compressed into learned visual latent embeddings, which condition a diffusion-based image generator.
Benchmarks
DreamBench AutoEval results across text-to-image generation, image editing alignment, and editing consistency.
Average accuracy (%) ↑
Average accuracy (%) ↑
Average accuracy (%) ↑
Gallery
Curated samples showcasing text-to-image generation, single-image editing, and multi-image composition.

An endless sea of sand stretches to the horizon. A robed figure walks atop the dune ridge, footprints trailing behind, sand kicked up by the wind.

Wrathful deity colossus standing amid raging flames. Glowing vajra, chains, and prayer flags. 8K ultra-clear, PBR materials, cinematic lighting, dark mythology aesthetic.

Green mountains, layered terraces, white-walled houses with dark roof tiles, wisps of cooking smoke, stone paths, murmuring streams, morning mist.

Lush green canopy, sunlight filtering through, moss-covered branches, hanging vines, morning mist, crystal-clear stream, ferns, mushroom clusters, firefly glow.

Abyssal vortex: a deep spiral water wall devouring ships, anti-gravity waterfalls flowing upward. 8K detail, cinematic lighting.

Night scene of Hong Kong's Mong Kok district. Dense neon signboards in vivid colors, deep street perspective, cinematic realism.

A fairytale cobblestone street with warm-toned houses, climbing vines on brick walls, blue sky, a cat lounging, and a bicycle leaning nearby.

Orange double-decker bus covered in dense pink flowers, number 10 clearly visible. Butterflies, flying petals, ultra-HD 8K 3D rendering.

An impasto emotional illustration at dusk — a girl in a white dress sits on a lake dock, her hair blown by the breeze, golden light reflecting off the water.

Meteorite impact, 8K ultra-clear. Lava eruption, shockwave expansion, earth crust cracking, city collapse, tsunami surge, cinematic catastrophe.
Pic 1
Pic 2
Pic 3
Pic 4
Pic 5
Pic 6
Pic 7
Pic 8
Pic 9
Pic 10
Pic 1
Pic 2
Pic 3
Pic 4
Pic 1
Pic 2
Pic 3
Pic 4
Pic 5
Pic 6
Pic 7
Pic 8
Pic 9
Pic 10
Head
Wings
Antennae
Legs
Tail
Pic 1
Pic 2
Pic 3
Pic 4
Pic 1
Pic 2
Pic 3
Pic 4
Team
ByteDance Seed
* Equal Contribution · † Project Lead
Citation