Z-Image-Engineer-V6, community fine-tune to boost Z-Image-Turbo
BennyDaBall released Z-Image-Engineer-V6, a 4B fine-tune of Alibaba's Z-Image-Turbo, to enhance prompts and act as a ComfyUI text encoder via SMART DoRA.
On the open-source front, Z-Image-Turbo established itself by late 2025: this 6-billion-parameter text-to-image model from Alibaba (Tongyi-MAI) generates in 8 steps in under a second, fits on 16 GB of VRAM, and ranks at the top of open-weights models on the Artificial Analysis Image Arena, notably for its bilingual text rendering.
Building on this model, a developer, BennyDaBall, has just released Z-Image-Engineer-V6, a 4-billion-parameter community fine-tune (Qwen encoder) with dual functionality. Locally (LM Studio, llama.cpp), it acts as a prompt enhancer: a minimal prompt becomes a structured description (composition, lighting, materials, depth), free of hollow phrases like “8k, masterpiece”. In ComfyUI, it replaces Z-Image-Turbo's original text encoder to produce different conditioning from the same seed.
The training relies on an in-house method, SMART DoRA, and its regularizers, which are intended to prevent repetitive loops. The author himself acknowledges: the tool improves adherence to composition, without guaranteeing a perfect seed every time.