TY - JOUR T1 - Assessing the Analytical and Clinical Validity of AI-Based Tumor-Infiltrating Lymphocyte Scoring in TNBC and the Interchangeability of Machine Learning Models A1 - Noah Williams A1 - Grace Collins A1 - Ethan Brooks A1 - Daniel Green JF - Interdisciplinary Research in Medical Sciences Specialty JO - Interdiscip Res Med Sci Spec SN - 3062-4401 Y1 - 2026 VL - 6 IS - 1 DO - 10.51847/eopncqTqQW SP - 271 EP - 283 N2 - Manual evaluation of tumor-infiltrating lymphocytes (TILs) by pathologists has revealed important predictive and prognostic roles in both early-stage and advanced triple-negative breast cancer (TNBC). Despite this, the approach continues to suffer from inconsistency. Artificial intelligence (AI) represents an effective way to minimize such inconsistency and support fully automated, impartial TILs measurement. The main hurdle to widespread clinical use is proving strong analytical and prognostic reliability. Ten different AI models were examined for their influence on TILs quantification, highlighting variations in analytical and prognostic performance. This included seven newly created models and three already validated ones. Testing occurred in a retrospective group for analytical aspects and a separate prospective group for prognostic aspects, using invasive disease-free survival (IDFS) as the outcome measure over a median 4-year follow-up period. Model development and analytical testing drew on diagnostic slides from 79 women treated for primary invasive TNBC at Yale School of Medicine between 2012 and 2016. Prognostic evaluation used an external group of 215 TNBC cases diagnosed in Sweden from 2010 to 2015. Marked differences emerged in analytical validity among the models and their training methods (Spearman’s r = 0.63–0.73, p < 0.001). Eight of the ten AI systems nonetheless displayed significant prognostic value for digital TILs scores in the independent external cohort. Hazard ratios were closely aligned and overlapping (HR = 0.40–0.47; p < 0.004) via Cox regression on the IDFS endpoint, even for models with more limited training. The consistent prognostic strength seen in most AI TIL models reflects the fundamental reliability of host anti-tumor immune response (captured through TILs) as a biomarker. Model-to-model differences, however, deserve attention. A shared, extensive, multi-institutional dataset is essential as a reference standard to enable fair comparisons and confirm dependability of diverse AI solutions ahead of routine clinical deployment. UR - https://galaxypub.co/article/assessing-the-analytical-and-clinical-validity-of-ai-based-tumor-infiltrating-lymphocyte-scoring-in-yeinbnxpcavr1jc ER -