Tasks

Public Example Tasks

The real 80-task benchmark corpus is private, since it's the held-out evaluation set. These 12 tasks are separate, hand-authored public examples in the exact same format (one easy, medium, and hard task per category), so anyone can see what a task looks like and run the tooling without access to the private set.

Public Examples

Every task in the public Hugging Face dataset, grouped by category and difficulty, loaded live (not bundled in this repository). Click a diagram to view it full size.