DT-RICH-26-2

Digital twin applications increasingly rely on large-scale, continuously operating computing infrastructure, yet current infrastructure operations remain largely reactive, fragmented, and manually intensive, particularly in system monitoring and issue diagnosis. We propose a Computing Infrastructure Digital Twin (CIDT) designed to provide what-now operational awareness while establishing a foundation for what-next forecasting to support future resource allocation, job scheduling, and optimized infrastructure usage. The core challenge addressed is the heavy reliance on manual labeling and rule-based monitoring, which limits scalability and responsiveness in modern AI/ML-intensive workloads. The key innovation is the automatic detection and classification of infrastructure issues using small and small-reasoning language models applied to real production system logs, enabling low-latency identification of failures beyond traditional approaches. This research will be tested on our own computing infrastructure and is intended to directly support more efficient and optimized computing allocation and scheduling for DT-RICH projects, serving as a practical step toward proactive and forecast-driven infrastructure management.