A Survey of Large Model Applications in Sports: From Multimodal Understanding to Intelligent Interaction

Published 2026-08-14 · AI Daily — AI-assisted deep research, methodology & disclosure

This article systematically reviews the current applications of large models in sports, focusing on how multimodal large language models are reshaping sports understanding, analysis, and interaction. It constructs a comprehensive landscape of tasks and applications covering various participant groups, analyzes existing sports-related datasets and benchmarks, and critically discusses current technical challenges and future research directions. The survey aims to lay a solid foundation for research and practical development of sports intelligence based on large models, while maintaining an open-source GitHub repository to support community collaboration. By integrating the latest advancements, the article reveals the immense potential of large models in enhancing sports health, cultural exchange, and economic growth, providing a clear roadmap for subsequent research.

Background and Context

The global sports industry is undergoing a significant transformation, driven by the increasing demand for sophisticated data analysis and personalized user experiences. Traditional methods of sports data processing have long been constrained by single-modal limitations, often failing to capture the full complexity and dynamic nature of athletic events. These conventional approaches typically isolate visual, auditory, or textual data, resulting in fragmented insights that do not reflect the holistic reality of a match or training session. The recent surge in global sports enthusiasm has highlighted the urgent need for systems that can integrate these diverse data streams to provide comprehensive support for health, cultural exchange, and economic growth.

This new survey, published on arXiv, addresses these historical gaps by providing a systematic review of large model applications specifically within the sports domain. The paper focuses on how multimodal large language models (MLLMs) are reshaping the landscape of sports understanding, tactical analysis, and human-machine interaction. By moving beyond isolated algorithmic improvements, the authors construct a comprehensive framework that maps the potential of these models across various participant groups. This includes athletes seeking performance optimization, coaches requiring tactical insights, and general audiences looking for enhanced viewing experiences. The study aims to fill a critical void in the literature, which has previously lacked a unified perspective on how large models can empower sports intelligence.

The core contribution of this research is the establishment of a clear roadmap for transitioning sports data analysis from a purely data-driven paradigm to one of cognitive intelligence. The authors emphasize that the value of large models lies not only in their technical capabilities but in their ability to adapt outputs to specific user roles. For instance, the same underlying model can generate detailed tactical breakdowns for a coach while providing real-time, engaging commentary for a spectator. This role-based adaptability represents a fundamental shift in how sports technology is designed, moving from static data recording to dynamic, intelligent interaction that serves diverse stakeholders within the sports ecosystem.

Deep Analysis

Technically, the survey constructs a multi-dimensional analysis framework that avoids the pitfalls of focusing on a single algorithm. Instead, it details the specific tasks and applications tailored to different participant groups, illustrating how large models adjust their output strategies based on user context. The paper provides a granular look at how these systems process information for athletes, coaches, and fans, demonstrating the versatility of multimodal architectures. This structured approach allows researchers and practitioners to understand the precise technical pathways required to deploy large models in real-world sports scenarios, moving from theoretical potential to practical implementation.

A critical component of the analysis involves a detailed examination of existing sports-related datasets and benchmarks. The authors identify significant shortcomings in current data resources, particularly regarding modal diversity, annotation quality, and scene coverage. They highlight that sports data is inherently unstructured, characterized by the temporal dependencies of video streams and the semantic ambiguity of textual descriptions. These characteristics pose unique challenges for model training, requiring strategies that can effectively handle the non-linear and complex nature of athletic events. The survey argues that current training methodologies must evolve to accommodate these specific data traits to ensure robust model performance.

Furthermore, the paper explores the mechanisms of multimodal fusion, detailing how visual, auditory, and textual information are integrated to deepen the model's understanding of sports events. The authors note that visual information plays a dominant role in action recognition, while textual data is indispensable for tactical interpretation. By dissecting these contributions, the survey reveals the specific logic and optimization directions for model architectures when processing dynamic sports scenes. This technical breakdown provides a clear view of how large models currently operate in sports contexts, identifying both the strengths of current fusion techniques and the areas requiring further refinement to handle the fast-paced nature of live sports.

Industry Impact

The publication of this survey has profound implications for both the open-source community and industrial applications. By maintaining an open-source GitHub repository, the authors have significantly lowered the barrier to entry for research in sports large models. This initiative fosters knowledge sharing and collaboration, allowing researchers and developers worldwide to build upon the established framework. The availability of a centralized resource for datasets, benchmarks, and code accelerates the pace of innovation, enabling the community to collectively address the technical challenges identified in the study.

For the industrial sector, the survey provides a theoretical and technical foundation for developing a range of intelligent sports products. These include automated match analysis tools, smart sports assistants, and personalized fitness recommendation systems. The insights gained from the study enable companies to design systems that move beyond passive data recording to active, intelligent analysis. This shift promises to enhance user experience and drive commercial value by offering deeper, more actionable insights to both professional teams and individual consumers. The potential for improving sports health and facilitating cultural exchange through these technologies further underscores the industry's interest in adopting these advanced models.

The survey also highlights the economic potential of integrating large models into the sports industry. By enabling more efficient analysis and personalized engagement, these technologies can unlock new revenue streams and improve operational efficiency. The ability to provide real-time, context-aware insights can transform how sports organizations interact with their audiences, creating more immersive and valuable experiences. This economic impact is complemented by the social benefits of promoting physical health and cultural connection, positioning sports intelligence as a key driver of broader societal value.

Outlook

Looking ahead, the survey identifies several critical directions for future research and development. One major focus is the need for more challenging and standardized benchmarks to accurately assess the true capabilities of large models in sports. Current evaluations often lack unified standards, leading to biased comparisons between different models. The authors advocate for the creation of new datasets that better reflect the complexity of real-world sports scenarios, ensuring that future models are rigorously tested against diverse and demanding tasks.

Additionally, the paper emphasizes the importance of addressing technical bottlenecks in fine-grained action recognition and long-term event reasoning. Current models still exhibit significant gaps in these areas, which are crucial for providing comprehensive tactical analysis. Future research should focus on improving the temporal reasoning capabilities of large models, enabling them to understand and predict complex sequences of events over extended periods. This will be essential for developing systems that can provide strategic insights to coaches and analysts.

Finally, the survey points to the need for advancements in privacy protection, model lightweighting, and cross-cultural adaptability. As sports intelligence becomes more integrated into daily life, ensuring the privacy of user data and reducing the computational cost of models will be critical for widespread adoption. Furthermore, developing models that can adapt to different cultural contexts and sports traditions will be key to achieving global applicability. By addressing these challenges, the sports intelligence ecosystem can mature, leading to more robust, efficient, and inclusive applications of large models in the sports domain.

Sources