ByteDance Seed

MMCORE

MultiModal COnnection with Representation Aligned Latent Embeddings — a unified framework for multimodal image generation and editing that bridges VLM reasoning with diffusion synthesis.

Bridging Understanding
and Generation

MMCORE leverages a pre-trained Vision-Language Model (VLM) to predict semantic visual embeddings via learnable query tokens, which serve as conditioning signals for a diffusion model. This streamlined design effectively transfers the rich understanding and reasoning capabilities of VLMs into the visual generation process, reducing computational overhead while maintaining high-fidelity synthesis.

  • Dual-Pathway Conditioning

    Visual query tokens capture global semantics while text embeddings preserve fine-grained lexical details.

  • Semantic Visual Alignment

    Distillation from frozen ViT encoders provides stable supervision for learning robust visual tokens.

  • Efficient Training

    Minimal overhead through efficient alignment of pretrained models, avoiding costly end-to-end retraining.

  • Interleaved Generation

    Scales to 10+ input images with fine-grained localized control for complex multi-image editing.

How MMCORE Works

Multimodal information from a VLM is compressed into learned visual latent embeddings, which condition a diffusion-based image generator.

MMCORE Architecture

Consistent Improvement over Seedream 4.0

DreamBench AutoEval results across text-to-image generation, image editing alignment, and editing consistency.

Text-to-Image Alignment

Average accuracy (%) ↑

MMCORE
84.4
GPT-Image-1
80.7
Seedream 4.0
78.2
Gemini 2.5
69.2
Qwen Image
59.1
Flux Kontext
57.7

Image Editing Alignment

Average accuracy (%) ↑

MMCORE
81.2
GPT-Image-1
79.9
Seedream 4.0
79.6
Gemini 2.5
75.8
Flux Kontext
58.7
Qwen Image
53.4

Editing Consistency

Average accuracy (%) ↑

MMCORE
70.6
Seedream 4.0
68.9
Gemini 2.5
68.0
Flux Kontext
54.8
Qwen Image
47.9
GPT-Image-1
42.4

Authors

ByteDance Seed

Zijie Li* Yichun Shi* Jingxiang Sun* Ye Wang Yixuan Huang Zhiyao Guo Xiaochen Lian Peihao Zhu Yu Tian Zhonghua Zhai† Peng Wang*†

* Equal Contribution  ·  † Project Lead

BibTeX

BibTeX
@article{li2026mmcore, title = {MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings}, author = {Li, Zijie and Shi, Yichun and Sun, Jingxiang and Wang, Ye and Huang, Yixuan and Guo, Zhiyao and Lian, Xiaochen and Zhu, Peihao and Tian, Yu and Zhai, Zhonghua and Wang, Peng}, journal = {arXiv preprint}, year = {2026} }