Large language models are changing enterprise analytics at two connected but historically separate layers. At the access layer, they translate business questions into queries, clarify intent, and explain results. At the delivery layer, they retrieve engineering context, configure workflows, generate transformation code, submit jobs, diagnose failures, and revise artifacts. This review argues that treating either layer as a stand-alone generation problem understates the central systems challenge: every useful action must remain grounded in schemas, dependencies, policies, and observable execution. SiriusBI and SiriusDeliver illustrate the emerging end-to-end design space. The first organizes conversational business intelligence around coordinated modules, multi-round clarification, and data-conditioned SQL-generation strategies; the second organizes warehouse delivery around hierarchical skill orchestration, lifecycle-aware artifact control, and learning from execution traces. Read alongside work on text-to-SQL, retrieval, tool use, data quality, provenance, and production technical debt, these systems suggest a common architecture built from constrained intent resolution, typed actions, executable verification, and accountable handoffs. The review synthesizes that architecture, identifies where benchmark evidence and production evidence answer different questions, and proposes design principles for analytics agents that can increase speed without obscuring responsibility.
- Jiang, J., Xie, H., Yang, J., Shen, S., Wang, Z., Zheng, Y., ... & Jiang, J. (2024). Siriusbi: A comprehensive llm-powered solution for data analytics in business intelligence. arXiv preprint arXiv:2411.06102.
- Xie, H., Zhou, X., Yang, J., Shen, S., Wang, Z., Zheng, Y., ... & Jiang, J. (2026). SiriusDeliver: Automating Data Warehouse Delivery at Tencent. arXiv preprint arXiv:2608.09185.
- Buneman, P., Khanna, S., & Tan, W. C. (2001). Why and where: A characterization of data provenance. In Database Theory - ICDT 2001 (pp. 316-330). Springer. https://doi.org/10.1007/3-540-44503-X_20 DOI
- Guo, J., Zhan, Z., Gao, Y., Xiao, Y., Lou, J. G., Liu, T., & Zhang, D. (2019). Towards complex text-to-SQL in cross-domain database with intermediate representation. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4524-4535). https://doi.org/10.18653/v1/P19-1444 DOI
- Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33.
- Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. Proceedings of the VLDB Endowment, 11(12), 1781-1794. https://doi.org/10.14778/3229863.3229867 DOI
- Schick, T., Dwivedi-Yu, J., Dessi, R., Raileanu, R., Lomeli, M., Hambro, E., et al. (2023). Toolformer: Language models can teach themselves to use tools. Advances in Neural Information Processing Systems, 36. https://doi.org/10.52202/075280-2997 DOI
- Scholak, T., Schucher, N., & Bahdanau, D. (2021). PICARD: Parsing incrementally for constrained auto-regressive decoding from language models. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (pp. 9895-9901). https://doi.org/10.18653/v1/2021.emnlp-main.779 DOI
- Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., et al. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28, 2503-2511.
- Wang, B., Shin, R., Liu, X., Polozov, O., & Richardson, M. (2020). RAT-SQL: Relation-aware schema encoding and linking for text-to-SQL parsers. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 7567-7578). https://doi.org/10.18653/v1/2020.acl-main.677 DOI
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2023). ReAct: Synergizing reasoning and acting in language models. In International Conference on Learning Representations.
- Yu, T., Zhang, R., Yang, K., Yasunaga, M., Wang, D., Li, Z., et al. (2018). Spider: A large-scale human-labeled dataset for complex and cross-domain semantic parsing and text-to-SQL task. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing (pp. 3911-3921). https://doi.org/10.18653/v1/D18-1425 DOI
- Journal
- AI Frontiers in Science and Society
- Volume
- 1 (2026)
- Article number
- osm20260004
- License
- CC BY 4.0
