Checking for non-preferred file/folder path names (may take a long time depending on the number of files/folders) ...

Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru


A newer version of this resource http://www.hydroshare.org/resource/1dfbdbe97e6d48d7b307591868ca70c4 is available that replaces this version.
Authors:
Owners: This resource does not have an owner who is an active HydroShare user. Contact CUAHSI (help@cuahsi.org) for information on this resource.
Type: Resource
Storage: The size of this resource is 6.3 GB
Created: Aug 11, 2025 at 10:11 p.m. (UTC)
Last updated: Jul 08, 2026 at 7:41 p.m. (UTC)
Citation: See how to cite this resource
Sharing Status: Public
Views: 1640
Downloads: 332
+1 Votes: 1 other +1 this
Comments: 1 comment

Abstract

We present the first comprehensive, open-access groundwater database for this region, developed by applying a novel workflow that integrates Optical Character Recognition (OCR) with Large Language Models (LLMs). This methodology systematically extracted and unified data from 3,675 previously unstructured 'gray literature' sources, including government reports and academic theses, spanning the period 1966–2024. The resulting database consolidates 4,775 multi-parameter data points—encompassing variables such as groundwater level, well elevation, water level (m a.s.l.), hydraulic conductivity, aquifer thickness, transmissivity, storage coefficient, specific yield, methodology used, and lithology, among others—and is augmented with 931 depth-to-water records from Peru's National Water Resources Information System (SNIRH). A rigorous three-phase quality assurance protocol, comprising 100% manual verification, statistical outlier screening, and spatial validation, was applied to ensure database accuracy. Designed to overcome data fragmentation, it enables trend analysis, groundwater modeling, and sustainable yield assessments. Publicly available via HydroShare with code on GitHub.

------------------------------------------

📌 Note on Dataset Versions:

- The file 'arequipa_final_gw_database.cvs' includes all the above data PLUS additional records sourced from the public database at [snirh.ana.gob.pe](https://snirh.ana.gob.pe), specifically for the following districts:
Apurímac, Arequipa, Ayacucho, Cusco, Ica, Moquegua, Puno, and Tacna.

- The file 'arequipa_final_gw_database_per_basin.xlsx' includes all of the aforementioned data, along with additional records sourced from the public database at snirh.ana.gob.pe. To facilitate easier classification and analysis, the workbook is organized so that each sheet corresponds to a specific basin.

➤ This extended version is intended for broader analysis and cross-validation with official national hydrological records.

Subject Keywords

Coverage

Spatial

Coordinate System/Geographic Projection:
WGS 84 EPSG:4326
Coordinate Units:
Decimal degrees
Place/Area Name:
Arequipa
North Latitude
-14.8800°
East Longitude
-71.3000°
South Latitude
-16.5000°
West Longitude
-73.3000°

Temporal

Start Date:
End Date:

Content

README.txt

--- START OF FILE README.txt ---

README: Regional Groundwater Database for Arequipa and Surroundings

Dataset Title: Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru
Authors: Héctor L. Venegas-Quiñones, Madeleine Guillen, et al.
Repository Platform: HydroShare
Date of Last Update: November 2025

1. OVERVIEW
-----------------
This repository contains the first comprehensive, open-access groundwater database for the Arequipa region and surrounding areas in Southern Peru. The dataset consolidates 4,775 multi-parameter data points encompassing depth to water, aquifer thickness, transmissivity, and storage coefficient, and is augmented with 931 depth-to-water records from Peru's National Water Resources Information System (SNIRH) spanning the period 1966–2024.

The data was derived from two primary streams:
1. "Gray Literature" Mining: 3,675 hydrogeological records (hydraulic conductivity, transmissivity, storage coefficient, lithology, etc.) extracted from 606 technical reports and academic theses using a semi-automated AI workflow (OCR + LLM) with 100% manual verification.
2. National Databases: 931 depth-to-water records sourced from Peru's National Water Resources Information System (SNIRH).

2. FILE STRUCTURE AND ACCESS
-----------------
IMPORTANT: Due to the comprehensive nature of the dataset, the database is compressed into a multi-part 7-Zip archive. These files contain the full raw data extraction in Excel format.

You must download ALL parts to access the data.

Files in this Repository:
- database.7z.001: Part 1 of the compressed archive.
- database.7z.002: Part 2 of the compressed archive.
- README.txt: This technical documentation file.

Instructions to Open:
1. Download both "database.7z.001" and "database.7z.002" to the same folder on your computer.
2. You need software capable of handling split archives (e.g., 7-Zip for Windows, Keka for macOS).
3. Right-click on the first file ("database.7z.001") and select "Extract Here".
4. The software will automatically combine Part 1 and Part 2.
5. This will extract the main Master File: 
   "arequipa_final_gw_database.csv"

3. DATA DICTIONARY (VARIABLE TRANSLATION)
-----------------
The column headers in the Excel file are in Spanish. The tables below provide the English translation and technical description.

A. SOURCE IDENTIFICATION (TRACEABILITY)
Spanish Header      | English Translation    | Description
------------------- | ---------------------- | --------------------------------------------------------
ID_Registro         | Record ID              | Unique identifier.
Codigo_Archivo      | File Code              | Internal code for the source PDF (e.g., ANA0000209).
Fuente_Completa     | Full Source            | Name of the institution or thesis repository.
Titulo_Documento    | Document Title         | Exact title of the original report/thesis.
Autor_es            | Author(s)              | Authors of the study.
Pagina_Referencia   | Page Reference         | Specific page(s) in the PDF where data was found.

B. SPATIAL & TEMPORAL
Spanish Header      | English Translation    | Description
------------------- | ---------------------- | --------------------------------------------------------
Fecha_Medicion      | Measurement Date       | YYYY-MM-DD.
Latitud_WGS84       | Latitude (WGS84)       | Standardized decimal latitude.
Longitud_WGS84      | Longitude (WGS84)      | Standardized decimal longitude.
Cota_msnm           | Elevation              | Meters above sea level.
CUENCA              | Basin                  | Hydrographic unit name.

C. HYDROGEOLOGICAL PARAMETERS
Spanish Header               | English Translation          | Unit
---------------------------- | ---------------------------- | --------
Nivel_Freatico               | Water Table Depth            | Meters (m)
Espesor_Saturado             | Saturated Thickness          | Meters (m)
Conductividad_Hidraulica_K   | Hydraulic Conductivity (K)   | See Unidad_K
Transmisividad_T             | Transmissivity (T)           | See Unidad_T
Coeficiente_Almacenamiento_S | Storage Coefficient (S)      | Dimensionless
Litologia                    | Lithology                    | Text

4. EXAMPLE: HOW TO TRACE DATA TO THE SOURCE
-----------------
To verify the origin of any data point in the "database.7z" files, use the 'Titulo_Documento' and 'Pagina_Referencia' columns.

Example (Row 2 in Excel):
- ID_Registro: 1
- Parameter: Transmisividad_T (Not listed in example row, but T would be here if measured) or Nivel_Freatico (176.05 masl).
- Source File: ANA0000209_1_compressed.pdf
- Document Title: "INVENTARIO Y EVALUACIÓN DE LAS FUENTES DE AGUA SUBTERRÁNEA EN EL VALLE DE ACARI"
- Author: MINISTERIO DE AGRICULTURA
- Page Reference: "54, 115"

Interpretation: The user can verify this record by finding the report "INVENTARIO... VALLE DE ACARI" (1981) and looking at pages 54 and 115, where this specific well data (ID 04/03/12-1) appears.

5. METHODOLOGY
-----------------
Data was extracted using a "Human-in-the-Loop" AI workflow:
1. OCR (Tesseract) converted scanned PDFs to text.
2. LLM (Gemini 2.5 Pro) parsed the text to extract hydrogeological parameters.
3. Validation: 100% of records were manually cross-checked by the authors against the source PDFs to ensure accuracy.

6. CITATION
-----------------
When using this dataset, please cite:
Venegas-Quiñones, H. L., et al. (2025). Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru. HydroShare.
--- END OF FILE README.txt ---

Related Resources

This resource has been replaced by a newer version Venegas-Quiñones, H. L., Guillen, M., Garcia-Chevesich, P. A., Uhle, B., González, E., Ticona, J., Díaz, J., Zea, J., Alejo, F., McCray, J. (2026). Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru, HydroShare, http://www.hydroshare.org/resource/1dfbdbe97e6d48d7b307591868ca70c4

Credits

Funding Agencies

This resource was created using funding from the following sources:
Agency Name Award Title Award Number
The Center for Mining Sustainability None 470266

How to Cite

Venegas-Quiñones, H. L., Guillen, M., Garcia-Chevesich, P. A., Uhle, B., González, E., Ticona, J., Díaz, J., Zea, J., Alejo, F., McCray, J. (2026). Regional Groundwater Database from Unstructured Sources Based on an Integrated Optical Character Recognition and Large Language Model Workflow for Arequipa and Surroundings in Southern Peru, HydroShare, http://www.hydroshare.org/resource/6664bf51199f461eb8bbad22edf78794

This resource is shared under the Creative Commons Attribution CC BY.

http://creativecommons.org/licenses/by/4.0/
CC-BY

Comments

New Comment

required