Evaluation practices for language-model analytics systems are fragmented. Text-to-SQL research offers controlled benchmarks and reproducible measures, while production reports emphasize completion rates, time savings, adoption, and user experience. Neither perspective alone establishes whether an analytics agent is dependable. This evidence synthesis uses SiriusBI and SiriusDeliver as connected cases: SiriusBI reports a modular conversational BI system with multi-round clarification and alternative SQL-generation strategies; SiriusDeliver reports an agent for warehouse task delivery with hierarchical skills, artifact lifecycle control, and trace-driven evolution. The article develops an evaluation matrix spanning semantic intent, executable artifacts, workflow outcomes, human work, governance, and longitudinal effects. It examines the limitations of exact match, aggregate success, autonomous-submission rates, and self-correction claims. Research on Spider, CoSQL, RAT-SQL, PICARD, cross-domain benchmark analysis, data-quality verification, production-readiness testing, human-AI interaction, data cascades, and model reporting supports a central conclusion: metrics should follow the system boundary and the consequences of error. A mature evaluation program combines offline replay, execution-based tests, human studies, staged production experiments, incident analysis, and post-deployment monitoring, with results stratified by risk and task difficulty.
- Jiang, J., Xie, H., Yang, J., Shen, S., Wang, Z., Zheng, Y., ... & Jiang, J. (2024). Siriusbi: A comprehensive llm-powered solution for data analytics in business intelligence. arXiv preprint arXiv:2411.06102.
- Xie, H., Zhou, X., Yang, J., Shen, S., Wang, Z., Zheng, Y., ... & Jiang, J. (2026). SiriusDeliver: Automating Data Warehouse Delivery at Tencent. arXiv preprint arXiv:2608.09185.
- Amershi, S., Weld, D., Vorvoreanu, M., Fourney, A., Nushi, B., Collisson, P., et al. (2019). Guidelines for human-AI interaction. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (pp. 1-13). https://doi.org/10.1145/3290605.3300233 DOI
- Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In 2017 IEEE International Conference on Big Data (pp. 1123-1132). https://doi.org/10.1109/BigData.2017.8258038 DOI
- Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daume III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86-92. https://doi.org/10.1145/3458723 DOI
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large language models cannot self-correct reasoning yet. In International Conference on Learning Representations. https://openreview.net/forum?id=IkmD3fKBPQ
- Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., et al. (2019). Model cards for model reporting. In Proceedings of the Conference on Fairness, Accountability, and Transparency (pp. 220-229). https://doi.org/10.1145/3287560.3287596 DOI
- Pourreza, M., & Rafiei, D. (2023). Evaluating cross-domain text-to-SQL models and benchmarks. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 1601-1614). https://aclanthology.org/2023.emnlp-main.99/
- Sambasivan, N., Kapania, S., Highfill, H., Akrong, D., Paritosh, P., & Aroyo, L. M. (2021). Everyone wants to do the model work, not the data work: Data cascades in high-stakes AI. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (pp. 1-15). https://doi.org/10.1145/3411764.3445518 DOI
- Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. Proceedings of the VLDB Endowment, 11(12), 1781-1794. https://doi.org/10.14778/3229863.3229867 DOI
- Scholak, T., Schucher, N., & Bahdanau, D. (2021). PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 9895-9901). https://doi.org/10.18653/v1/2021.emnlp-main.779 DOI
- Wang, B., Shin, R., Liu, X., Polozov, O., & Richardson, M. (2020). RAT-SQL: Relation-aware schema encoding and linking for text-to-SQL parsers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7567-7578). https://doi.org/10.18653/v1/2020.acl-main.677 DOI
- Yu, T., Zhang, R., Er, H., Li, S., Xue, E., Pang, B., et al. (2019). CoSQL: A conversational text-to-SQL challenge towards cross-domain natural language interfaces to databases. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 1962-1979). https://doi.org/10.18653/v1/D19-1204 DOI
- Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., et al. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 3911-3921). https://doi.org/10.18653/v1/D18-1425 DOI
- Journal
- AI Frontiers in Science and Society
- Volume
- 1 (2026)
- Article number
- osm20260005
- License
- CC BY 4.0
