Model leaderboard

Every model we’ve run through the standard Glif tests, scored out of 10 by an LLM judge. Overall needs 3 judged tests to rank; pick a capability to rank on it alone.

Creative interpretation

Given a vague, abstract, or nonsense prompt, produces something intentional, aesthetic, and conceptually interesting rather than bland or incoherent.

  1. 1

    GPT Image 1.5

    8.6

  2. 2

    FLUX.2 Max Image Generator

    8.5

  3. 3

    Nano Banana Pro Image Generator

    8.5

  4. 4

    FLUX.2 Flex Image Generator

    8.3

  5. 5

    FLUX.2 Pro Image Generator

    8.3

  6. 6

    FLUX.2 Turbo Image Generator

    8.0

  7. 7

    FLUX.1 Dev Image Generator

    7.9

  8. 8

    Seedream V4.5

    7.5

  9. 9

    FLUX.1 Schnell Image Generator

    7.3

  10. 10

    Qwen Image Generator with LoRA

    6.8

  11. 11

    FLUX.2 Klein 9B Image Generator

    6.3