Checking for non-preferred file/folder path names (may take a long time depending on the number of files/folders) ...
This resource contains some files/folders that have non-preferred characters in their name. Show non-conforming files/folders.
This resource contains content types with files that need to be updated to match with metadata changes. Show content type files that need updating.
Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru
| Authors: |
|
|
|---|---|---|
| Owners: |
|
This resource does not have an owner who is an active HydroShare user. Contact CUAHSI (help@cuahsi.org) for information on this resource. |
| Type: | Resource | |
| Storage: | The size of this resource is 6.3 GB | |
| Created: | Jun 09, 2026 at 6:52 p.m. (UTC) | |
| Last updated: | Jun 17, 2026 at 6:27 p.m. (UTC) | |
| Citation: | See how to cite this resource |
| Sharing Status: | Public |
|---|---|
| Views: | 175 |
| Downloads: | 29 |
| +1 Votes: | Be the first one to this. |
| Comments: | No comments (yet) |
Abstract
We present the first comprehensive, open-access groundwater database for this region, developed by applying a novel workflow that integrates Optical Character Recognition (OCR) with Large Language Models (LLMs). This methodology systematically extracted and unified data from 3,675 previously unstructured 'gray literature' sources, including government reports and academic theses, spanning the period 1966–2024. The resulting database consolidates 4,775 multi-parameter data points—encompassing variables such as groundwater level, well elevation, water level (m a.s.l.), hydraulic conductivity, aquifer thickness, transmissivity, storage coefficient, specific yield, methodology used, and lithology, among others—and is augmented with 931 depth-to-water records from Peru's National Water Resources Information System (SNIRH). A rigorous three-phase quality assurance protocol, comprising 100% manual verification, statistical outlier screening, and spatial validation, was applied to ensure database accuracy. Designed to overcome data fragmentation, it enables trend analysis, groundwater modeling, and sustainable yield assessments. Publicly available via HydroShare with code on GitHub.
------------------------------------------
📌 Note on Dataset Versions:
- The file 'arequipa_final_gw_database.cvs' includes all the above data PLUS additional records sourced from the public database at [snirh.ana.gob.pe](https://snirh.ana.gob.pe), specifically for the following districts:
Apurímac, Arequipa, Ayacucho, Cusco, Ica, Moquegua, Puno, and Tacna.
- The file 'arequipa_final_gw_database_per_basin.xlsx' includes all of the aforementioned data, along with additional records sourced from the public database at snirh.ana.gob.pe. To facilitate easier classification and analysis, the workbook is organized so that each sheet corresponds to a specific basin.
➤ This extended version is intended for broader analysis and cross-validation with official national hydrological records.
Subject Keywords
Coverage
Spatial
Temporal
| Start Date: | |
|---|---|
| End Date: |
Content
README.txt
--- START OF FILE README.txt ---
README: Regional Groundwater Database for Arequipa and Surroundings
Dataset Title: Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru
Authors: Héctor L. Venegas-Quiñones, Madeleine Guillen, et al.
Repository Platform: HydroShare
Date of Last Update: November 2025
1. OVERVIEW
-----------------
This repository contains the first comprehensive, open-access groundwater database for the Arequipa region and surrounding areas in Southern Peru. The dataset consolidates 4,775 multi-parameter data points encompassing depth to water, aquifer thickness, transmissivity, and storage coefficient, and is augmented with 931 depth-to-water records from Peru's National Water Resources Information System (SNIRH) spanning the period 1966–2024.
The data was derived from two primary streams:
1. "Gray Literature" Mining: 3,675 hydrogeological records (hydraulic conductivity, transmissivity, storage coefficient, lithology, etc.) extracted from 606 technical reports and academic theses using a semi-automated AI workflow (OCR + LLM) with 100% manual verification.
2. National Databases: 931 depth-to-water records sourced from Peru's National Water Resources Information System (SNIRH).
2. FILE STRUCTURE AND ACCESS
-----------------
IMPORTANT: Due to the comprehensive nature of the dataset, the database is compressed into a multi-part 7-Zip archive. These files contain the full raw data extraction in Excel format.
You must download ALL parts to access the data.
Files in this Repository:
- database.7z.001: Part 1 of the compressed archive.
- database.7z.002: Part 2 of the compressed archive.
- README.txt: This technical documentation file.
Instructions to Open:
1. Download both "database.7z.001" and "database.7z.002" to the same folder on your computer.
2. You need software capable of handling split archives (e.g., 7-Zip for Windows, Keka for macOS).
3. Right-click on the first file ("database.7z.001") and select "Extract Here".
4. The software will automatically combine Part 1 and Part 2.
5. This will extract the main Master File:
"arequipa_final_gw_database.csv"
3. DATA DICTIONARY (VARIABLE TRANSLATION)
-----------------
The column headers in the Excel file are in Spanish. The tables below provide the English translation and technical description.
A. SOURCE IDENTIFICATION (TRACEABILITY)
Spanish Header | English Translation | Description
------------------- | ---------------------- | --------------------------------------------------------
ID_Registro | Record ID | Unique identifier.
Codigo_Archivo | File Code | Internal code for the source PDF (e.g., ANA0000209).
Fuente_Completa | Full Source | Name of the institution or thesis repository.
Titulo_Documento | Document Title | Exact title of the original report/thesis.
Autor_es | Author(s) | Authors of the study.
Pagina_Referencia | Page Reference | Specific page(s) in the PDF where data was found.
B. SPATIAL & TEMPORAL
Spanish Header | English Translation | Description
------------------- | ---------------------- | --------------------------------------------------------
Fecha_Medicion | Measurement Date | YYYY-MM-DD.
Latitud_WGS84 | Latitude (WGS84) | Standardized decimal latitude.
Longitud_WGS84 | Longitude (WGS84) | Standardized decimal longitude.
Cota_msnm | Elevation | Meters above sea level.
CUENCA | Basin | Hydrographic unit name.
C. HYDROGEOLOGICAL PARAMETERS
Spanish Header | English Translation | Unit
---------------------------- | ---------------------------- | --------
Nivel_Freatico | Water Table Depth | Meters (m)
Espesor_Saturado | Saturated Thickness | Meters (m)
Conductividad_Hidraulica_K | Hydraulic Conductivity (K) | See Unidad_K
Transmisividad_T | Transmissivity (T) | See Unidad_T
Coeficiente_Almacenamiento_S | Storage Coefficient (S) | Dimensionless
Litologia | Lithology | Text
4. EXAMPLE: HOW TO TRACE DATA TO THE SOURCE
-----------------
To verify the origin of any data point in the "database.7z" files, use the 'Titulo_Documento' and 'Pagina_Referencia' columns.
Example (Row 2 in Excel):
- ID_Registro: 1
- Parameter: Transmisividad_T (Not listed in example row, but T would be here if measured) or Nivel_Freatico (176.05 masl).
- Source File: ANA0000209_1_compressed.pdf
- Document Title: "INVENTARIO Y EVALUACIÓN DE LAS FUENTES DE AGUA SUBTERRÁNEA EN EL VALLE DE ACARI"
- Author: MINISTERIO DE AGRICULTURA
- Page Reference: "54, 115"
Interpretation: The user can verify this record by finding the report "INVENTARIO... VALLE DE ACARI" (1981) and looking at pages 54 and 115, where this specific well data (ID 04/03/12-1) appears.
5. METHODOLOGY
-----------------
Data was extracted using a "Human-in-the-Loop" AI workflow:
1. OCR (Tesseract) converted scanned PDFs to text.
2. LLM (Gemini 2.5 Pro) parsed the text to extract hydrogeological parameters.
3. Validation: 100% of records were manually cross-checked by the authors against the source PDFs to ensure accuracy.
6. CITATION
-----------------
When using this dataset, please cite:
Venegas-Quiñones, H. L., et al. (2025). Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru. HydroShare.
--- END OF FILE README.txt ---
Related Resources
| This resource updates and replaces a previous version | Venegas-Quiñones, H. L., M. Guillen, P. A. Garcia-Chevesich, B. Uhle, E. González, J. Ticona, J. Díaz, J. Zea, F. Alejo, J. McCray (2026). Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru, HydroShare, http://www.hydroshare.org/resource/6664bf51199f461eb8bbad22edf78794 |
Credits
Funding Agencies
This resource was created using funding from the following sources:
| Agency Name | Award Title | Award Number |
|---|---|---|
| The Center for Mining Sustainability | None | 470266 |
How to Cite
This resource is shared under the Creative Commons Attribution CC BY.
http://creativecommons.org/licenses/by/4.0/
Comments
There are currently no comments
New Comment