Abstract
Recent advances in large language models (LLMs) have driven remarkable progress across diverse NLP benchmarks. However, the application of these models to sophisticated domains such as wireless communications raises new challenges. This paper addresses the evaluation gap for LLMs in telecommunications by introducing Last Dance of Telecommunications (LDOT) – a comprehensive benchmark suite for telecom-related tasks. LDOT encompasses a broad range of problem categories, including conceptual telecom questions, mathematical and logical reasoning problems, and complex network optimization scenarios, specifically designed to challenging LLMs’ high-level reasoning and domain-specific knowledge. Using LDOT, we rigorously assess state-of-the-art (SOTA) LLMs (both closed-source and open-source) on domain-specific tasks. Our results reveal that while general-purpose LLMs exhibit strong performance on basic telecom knowledge questions, they struggle with reasoning-intensive wireless problems. Notably, certain multi-step optimization and planning tasks in LDOT remain unsolved by even the best models, exposing performance gaps that are not apparent from existing saturated benchmarks. We provide a detailed failure analysis to pinpoint whether these limitations arise from insufficient telecom-specific knowledge or from inadequate reasoning capabilities. The LDOT dataset and our evaluation findings aim to facilitate the development of more robust domain-adapted LLMs for next-generation wireless communications.
| Original language | British English |
|---|---|
| Pages (from-to) | 2365-2377 |
| Number of pages | 13 |
| Journal | IEEE Journal on Selected Areas in Communications |
| Volume | 44 |
| DOIs | |
| State | Published - 2026 |
Keywords
- dataset
- evaluation
- Large language models (LLMs)
- multi-agent
- telecommunications
Fingerprint
Dive into the research topics of 'Go Gentle Into the Final Dance: A Benchmark for Evaluating LLMs in Telecom'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver