Tar by ByteDance
verifiedFreecsuhan.com
Open research model (ByteDance/CUHK) unifying image understanding and generation via text-aligned tokens.
What it does
Tar is an open research project presenting a multimodal LLM that unifies visual understanding and generation. It uses a Text-Aligned Tokenizer to turn images into discrete tokens drawn from an LLM vocabulary, letting a single model both interpret and create images. It was published at NeurIPS 2025 with public code and demos.
Core features
Text-Aligned Tokenizer (TA-Tok)
Unified understanding and generation
Scale-adaptive encoding/decoding
Autoregressive and diffusion de-tokenizers
Built on a Qwen2 backbone
Open code and demos
Best for
→Image captioning and visual Q&A
→Text-to-image generation
→Multimodal AI research
→Benchmarking unified models