Bugs in llm_annotator.ipynb prompt: swapped 3.2/3.3 labels and stray 1.6/2.7 codes in few-shot example
MAD_human_labelled_dataset.json spans three taxonomy versions with renumbered codes — a simpler explanation for #13
MAD_full_dataset.json: remaining cases of #6 — 510/1242 records share replicated annotation sequences
MAD_human_labelled_dataset.json: 19 records share only 8 distinct annotation blocks
multi-agent-systems-failure-taxonomy/MAST
Empirical results: F1 0.815 on mode 3.3 (deterministic detector, Fleiss κ=1.000), MAD per-MAS breakdown, methodology notes
Published failure_modes in MAD_human_labelled_dataset.json are inconsistent with the dataset's own embedded messages.note.options and with outputs from the repo's llm_annotator.ipynb
Questions about the MAST dataset and evaluation methodology
Any end-to-end example?
Multi-Agent System Code Built on AppWorld