Skip to content

About

FinAuditing: a taxonomy-structured, multi-document auditing benchmark for LLMs (arXiv:2510.08886)

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Repository files navigation

FinAuditing

A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs

arXiv SIGIR 2026 Resource Track Hugging Face Collection License: MIT

Paper · Data · Evaluation Framework


Overview

FinAuditing is a taxonomy-aligned, structure-aware benchmark built from real XBRL filings for evaluating whether LLMs can perform professional-grade financial auditing. It defines three tasks: FinSM (Financial Semantic Matching), FinRE (Financial Relationship Extraction), and FinMR (Financial Mathematical Reasoning), each grounded in an XBRL filing and the US-GAAP taxonomy. This repository provides a starter kit for local inference and the task-specific evaluation notebooks.

Getting Started

We provide two evaluation pathways, depending on whether you prefer lightweight local testing or full benchmark evaluation via our unified framework.

1. Local testing via the Starter Kit (recommended for quick experiments)

To test the benchmark with local code:

  • See the StartKit/ directory.
  • We provide two end-to-end pipelines:
    • pipeline-hf.ipynb for Hugging Face-based inference
    • pipeline-vllm.ipynb for vLLM-based inference
  • After inference, use the three task-specific evaluation scripts in the same directory to evaluate performance on FinSM, FinRE, and FinMR, respectively.

2. Full evaluation via the FinBen framework (recommended for benchmarking)

To evaluate models with our unified evaluation framework (FinBen fork):

Resources on Hugging Face

All FinAuditing data is collected in the FinAuditing Hugging Face collection.

Dataset Description
TheFinAI/en-finsm (formerly FinSM) Evaluation set for the FinSM subtask of the FinAuditing benchmark. This task follows the information retrieval paradigm: given a query describing a financial term that represents either currency or concentration of credit risk, an XBRL filing, and a US-GAAP taxonomy, the output is the set of mismatched US-GAAP tags after retrieval.
TheFinAI/en-finre (formerly FinRE) Evaluation set for the FinRE subtask of the FinAuditing benchmark. This is a relation extraction task: given two specific elements $e_1$ and $e_2$, an XBRL filing, and a US-GAAP taxonomy, the goal is to classify three relation error types.
TheFinAI/en-finmr (formerly FinMR) Evaluation set for the FinMR subtask of the FinAuditing benchmark. This is a mathematical reasoning task: given two questions $q_1$ and $q_2$, where $q_1$ concerns the extraction of a reported value and $q_2$ pertains to the calculation of the corresponding real value, an XBRL filing, and a US-GAAP taxonomy, the task is to extract the reported value for a given instance in the XBRL filing and to compute the numeric value for that instance, which is then used to verify whether the reported value is correct.
TheFinAI/en-finsm-sub (formerly FinSM_Sub) FinSM subset for ICAIF 2026.
TheFinAI/en-finre-sub (formerly FinRE_Sub) FinRE subset for ICAIF 2026.
TheFinAI/en-finmr-sub (formerly FinMR_Sub) FinMR subset for ICAIF 2026.

Citation

If you find our benchmark useful, please cite:

@misc{wang2025finauditingfinancialtaxonomystructuredmultidocument,
      title={FinAuditing: A Financial Taxonomy-Structured Multi-Document Benchmark for Evaluating LLMs}, 
      author={Yan Wang and Keyi Wang and Shanshan Yang and Jaisal Patel and Jeff Zhao and Fengran Mo and Xueqing Peng and Lingfei Qian and Jimin Huang and Guojun Xiong and Xiao-Yang Liu and Jian-Yun Nie},
      year={2025},
      eprint={2510.08886},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2510.08886}, 
}

License

The code in this repository is released under the MIT License. Datasets and models on Hugging Face keep their own licenses, stated on each card.


Built by The Fin AI · Hugging Face · GitHub

About

FinAuditing: a taxonomy-structured, multi-document auditing benchmark for LLMs (arXiv:2510.08886)

Topics

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages