2026
Zero-Shot Hierarchical Tagging of Short Technical Posts with Large Language Models
A case study evaluating a schema-constrained LLM on 80 technical X posts against a 155-node personal taxonomy. The LLM reaches hierarchical F1 0.622 versus 0.574 (TF–IDF) and 0.569 (MiniLM); paired bootstrap intervals include zero. All 180 held-out responses are valid on first attempt; mean pairwise Jaccard is 0.755.