Submitted by Remco Hendriks 14 MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes Continker 0 4