toolspool
Tar by ByteDance logo

Tar by ByteDance

verifiedFreecsuhan.com

Open research model (ByteDance/CUHK) unifying image understanding and generation via text-aligned tokens.

What it does

Tar is an open research project presenting a multimodal LLM that unifies visual understanding and generation. It uses a Text-Aligned Tokenizer to turn images into discrete tokens drawn from an LLM vocabulary, letting a single model both interpret and create images. It was published at NeurIPS 2025 with public code and demos.

Core features

Text-Aligned Tokenizer (TA-Tok)
Unified understanding and generation
Scale-adaptive encoding/decoding
Autoregressive and diffusion de-tokenizers
Built on a Qwen2 backbone
Open code and demos

Best for

Image captioning and visual Q&A
Text-to-image generation
Multimodal AI research
Benchmarking unified models