Institutional Investors

How LLMs are revolutionising institutional investment data extraction

Large Language Models (LLMs) are fundamentally changing investment data extraction by moving beyond rigid templates and basic text recognition. They can now read, understand context, and extract complex financial data from any unstructured document, much like a human analyst.

What was wrong with old data extraction?

Previous attempts at data extraction in finance relied on two flawed technologies:

  1. Basic OCR: Optical Character Recognition simply turns a PDF image into a text file. It doesn’t understand what the text means.
  2. Template-based extraction: This rigid approach required developers to define specific coordinates for each data point (e.g., “NAV is always on Page 3, line 5”). The entire system would break the moment a GP changed their report format.

How are LLMs different?

LLMs, the same technology powering models like Claude, Gemini and ChatGPT, work differently. They understand semantic meaning and context.

  • They understand context: An LLM knows that “Net Asset Value,” “NAV,” and “Partners’ Capital” might mean the same thing, depending on the context.
  • They are format-agnostic: They don’t rely on a fixed template. They can find the NAV whether it’s in a table, a footnote, or a narrative sentence.
  • They handle complexity: They can extract not just simple numbers but also complex relationships, such as the data from a multi-layered Schedule of Investments.

This is the same principle behind modern AI search, which retrieves information based on semantic meaning, not just keywords.

What new capabilities does this unlock for investors?

This technological leap unlocks capabilities that were previously impossible:

  • True automation: The ability to process 99% of documents from any GP without manual intervention.
  • Extracting qualitative data: LLMs can summarise management commentary, identify risk factors mentioned in narrative text, and classify the sentiment of a report.
  • Handling the “long tail”: They can process the ad-hoc, one-off, and non-standard documents that template-based systems could never handle.
  • Validation and audit: You can “ask” the model why it extracted a certain number, and it can point to the exact source text in the original document.

How is Scribe using this technology?

Scribe applies this technology to deliver an intelligent co-pilot purpose-built for private markets. By using Large Language Models to extract financial data accurately and at scale, we transform unstructured reporting into structured, analysable data, allowing your team to focus on the “Analyse” and “Invest” functions that drive returns.