Model leaderboard

Every model we’ve run through the standard Glif tests, scored out of 10 by an LLM judge. Overall needs 3 judged tests to rank; pick a capability to rank on it alone.

Localized editing

Changes only what was asked and leaves the rest of the image pixel-faithful. Penalize global drift, recrops, or unrequested changes.

  1. 1

    GPT Image 1.5

    6.8

  2. 2

    FLUX Kontext Image Editor

    3.0

  3. 3

    Qwen Image Generator with LoRA

    1.8