This document describes the external and internal data providers used in Climate Loop, their characteristics, limitations, and integration strategy, and how they support synthetic IoT data generation and autonomous forecasting.
Documentation: https://open-meteo.com/en/docs/historical-forecast-api
Open-Meteo provides free access to historical forecast and reanalysis-based weather data with hourly resolution. It aggregates multiple upstream numerical weather prediction (NWP) and reanalysis sources behind a unified API.
In Climate Loop, Open-Meteo is used as the primary structured hourly time-series backbone for:
- Continuous weather prediction (temperature, precipitation, wind, humidity)
- Feature engineering (lags, rolling windows, seasonality encoding)
- Synthetic disaster signal generation
- Multi-city benchmarking
Example parameters used:
latitude,longitudestart_date,end_datehourlyvariables:temperature_2mrelative_humidity_2mdew_point_2mapparent_temperatureprecipitation_probabilityprecipitationrain,showers,snowfallpressure_msl,surface_pressurevisibilitywind_speed_80m,wind_direction_80mwind_speed_180m,wind_direction_180mwind_gusts_10msoil_temperature_0cm,soil_temperature_18cmevapotranspiration,et0_fao_evapotranspirationweather_code
Temporal coverage used in the project:
- 2000-01-01 → 2026-02-22
- Hourly granularity
| Property | Value |
|---|---|
| Temporal Resolution | Hourly |
| Spatial Resolution | Model grid (~10–30 km depending on upstream model) |
| Data Type | Forecast + Reanalysis-backed |
| Access | REST API (JSON) |
| Cost | Free tier |
- Consistent schema across cities.
- High-resolution hourly data.
- Broad variable coverage (surface + atmospheric + soil).
- Easy automation and reproducibility.
- Not ground-truth station data (model-based).
- Grid-based interpolation may smooth extreme events.
- Forecast uncertainty is not always directly exposed.
- Changes in upstream models may introduce temporal distribution shifts.
Open-Meteo is not used solely for dataset standardization. It is a core historical data provider that enables long-horizon model training and cross-source validation.
Specifically, Open-Meteo data is:
- Used to build a multi-year historical climate dataset per city (2000–2026).
- Combined with Meteostat station observations to increase robustness.
- Leveraged to cross-validate model-based reanalysis data against real station measurements.
- Used to generate consistent feature spaces across cities.
- The primary training backbone for continuous forecasting models (LSTM, XGBoost, VAR).
The strategic objective is:
-
Avoid runtime dependency on third-party APIs.
Models are trained offline on historical official sources so predictions can be generated independently if APIs become unavailable. -
Improve realism through multi-source alignment.
By combining model-based reanalysis data (Open-Meteo) with station-observed measurements (Meteostat), the system reduces bias and improves representation of extreme events. -
Continuously improve model fidelity.
The architecture is designed to incorporate additional official climate sources when available.
If a new dataset improves signal quality or extreme-event realism, it can be integrated into the training pipeline.
Station ID Example: 83755
Documentation: https://meteostat.net/
Meteostat provides access to historical weather observations from physical meteorological stations worldwide.
In Climate Loop, Meteostat is used as a ground-truth validation layer to:
- Compare model-based Open-Meteo data against station observations.
- Evaluate bias and systematic deviation.
- Assess extreme event realism.
- Perform cross-source validation studies.
| Property | Value |
|---|---|
| Temporal Resolution | Daily / Hourly (station dependent) |
| Data Type | Observed station measurements |
| Spatial Resolution | Point-based (station coordinates) |
| Access | API / CSV exports |
| Cost | Free access tier |
- Real measured observations (not simulated).
- Better representation of localized extremes.
- Useful for calibration and bias correction.
- Missing data periods.
- Inconsistent variable availability across stations.
- Sensor drift or station relocation effects.
- Lower spatial representativeness compared to grid models.
Meteostat data is:
- Used for validation and descriptive statistical analysis.
- Compared against Open-Meteo in bias studies.
- Used in
compare_era5_meteo.ipynb.
Official CAP alerts are structured emergency notifications issued by authorized entities such as meteorological agencies, civil protection authorities, and emergency management organizations. They communicate imminent or ongoing hazards including severe storms, floods, heatwaves, wildfires, strong winds, and other public safety risks.
Unlike continuous meteorological time-series data, CAP alerts are event-driven declarations. They are triggered when predefined risk thresholds are reached or when authorities determine that public notification is necessary.
CAP (Common Alerting Protocol) is an international XML-based standard designed to ensure interoperability across emergency communication systems. It defines structured fields such as:
- event – type of hazard
- severity – level of risk
- urgency – time-criticality
- certainty – likelihood
- areaDesc – affected geographic region
- effective and expires times
- description – official text and recommended actions
Because CAP is machine-readable and schema-based, alerts can be automatically parsed, validated, and integrated into automated decision-support systems such as Climate Loop.
Official CAP alerts are typically published via government-managed RSS feeds or dedicated XML endpoints maintained by national meteorological and civil protection authorities.
In Climate Loop, the alerts ingestion process operates as follows:
- The system polls official RSS feeds published by agencies such as AEMET (Spain) or regional authorities.
- It extracts embedded CAP XML from .tar.gz packages or direct XML links.
- Structured fields are parsed (e.g., severity, urgency, geographic impact, timing).
- The alerts are filtered based on the user’s real-time location.
- Relevant alerts are stored as structured event objects in the platform.
These alerts are used directly in the application interface. When a user opens the app, Climate Loop checks whether any active alerts apply to their location. If relevant alerts exist, they are surfaced prominently in the UI.
Once an alert is identified as relevant, the platform’s Generative AI layer processes the raw institutional alert data to produce:
- Plain-language explanations
- Risk summaries adapted to local context
- Actionable protective guidance
- Multilingual support based on user preference
This ensures that official emergency declarations are not only displayed, but translated into clear, understandable, and practical guidance that helps users know what action to take.
Below is an example of an official RSS feed from AEMET (Agencia Estatal de Meteorología) for meteorological warnings in Andalucía. This example shows the format even when no active warnings exist. When warnings are active, the associated CAP files contain structured XML alert data.
<rss version="2.0">
<channel>
<title>AEMET - Avisos de fenómenos meteorológicos adversos - Andalucía</title>
<link>https://www.aemet.es/es/eltiempo/prediccion/avisos?k=and</link>
<description>
Avisos de fenómenos meteorológicos adversos según el Plan Meteoalerta
de la Agencia Estatal de Meteorología (AEMET) para Andalucía.
</description>
<language>es</language>
<copyright>https://www.aemet.es/es/nota_legal</copyright>
<pubDate>2026-02-23T22:50:01+00:00</pubDate>
<lastBuildDate>2026-02-23T22:50:01+00:00</lastBuildDate>
<image>
<url>https://www.aemet.es/imagenes/gif/logo_AEMET_web.gif</url>
</image>
<item>
<title>
Estado completo de avisos para Andalucía. No hay avisos para Andalucía
</title>
<description>
Fichero tar.gz que contiene todos los avisos para Andalucía de
23:50 23-02-2026 CET (UTC+1) a 23:59 26-02-2026 CET (UTC+1).
</description>
<link>
https://www.aemet.es/documentos_d/eltiempo/prediccion/avisos/cap/Z_CAP_C_LEMM_20260223225001_AFAC61.tar.gz
</link>
<guid>Z_CAP_C_LEMM_20260223225001_AFAC61.tar.gz</guid>
<pubDate>2026-02-23T22:50:01+00:00</pubDate>
</item>
</channel>
</rss>When active alerts are present, the linked CAP files contain structured XML describing:
- The hazard type
- Severity level
- Urgency and certainty
- Affected geographic regions
- Official protective recommendations
These alerts provide a critical event-driven signal that complements continuous forecasting models:
- They act as authoritative confirmation of extreme events.
- They trigger real-time contextual guidance generation.
- They align predictive risk modeling with institutional declarations.
- They feed into the Generative AI layer for clear, user-friendly explanations.
- They are delivered directly to users based on geographic relevance.
By integrating CAP alerts with forecasting data and community signals, Climate Loop ensures both technical accuracy and practical intelligibility during elevated risk situations.
Since no proprietary IoT device infrastructure is currently deployed, the project generates synthetic IoT-like datasets using historical climatological sources.
A physical proof of concept device was developed and tested to validate the feasibility of local environmental data collection. This prototype allowed us to determine which environmental variables could realistically be captured, at what frequency, and with what level of signal stability. Additional technical details about the hardware and a short demonstration video are provided in the Quick Note on IoT Hardware section below.
The insights obtained from this proof of concept directly informed the structure, feature selection, and schema design of the synthetic IoT datasets used in this phase.
The approach implemented is:
-
Historical Source Fusion
- Multi-year Open-Meteo data (grid-based reanalysis/forecast).
- Station-observed data from Meteostat.
- Alignment in timestamp, units, and feature schema.
-
Signal Calibration
- Station data is used to identify bias and variance differences.
- Statistical adjustments are applied where appropriate.
- Extreme events are preserved to avoid over-smoothing.
-
IoT Profile Simulation
- Hourly resolution is enforced.
- Sensor-like noise can be injected (controlled stochastic perturbations).
- Feature subsets mimic realistic IoT device constraints (e.g., limited sensor types).
-
Dataset Output
- Generated as
iot_hourly_{city}.csv. - Structured to match expected production IoT schema.
- Used for ML model training and pipeline validation.
- Generated as
The goal of this synthetic IoT generation layer is to:
- Develop forecasting models before real IoT deployment.
- Train forecasting models on realistic climatological distributions.
- Reduce dependency on external APIs at inference time.
- Provide a transition-ready architecture: when real IoT data becomes available, it can replace the synthetic layer without refactoring downstream ML components.
To validate the feasibility of local environmental data collection, a physical proof of concept device was assembled using:
- Arduino Uno R1
- BMP280 sensor for temperature and atmospheric pressure
- Rain droplet detection module
A short demonstration of this prototype can be viewed here:
https://www.youtube.com/shorts/rbx0PtpRxHs
The Arduino Uno R1 is an older model and is no longer widely commercialized, making it difficult to estimate its current price. However, newer Arduino-compatible boards typically cost under €30.
The additional components used in the prototype are low cost:
- BMP280 sensor: under €2
- Rain detection module: under €6
This demonstrates that a basic environmental sensing unit can be assembled at very low hardware cost.
For future iterations, the plan is to migrate to more modern and integrated microcontrollers such as the ESP32-CAM. This device includes built-in WiFi and a camera module, allowing automatic sky image capture for visual weather monitoring. The ESP32-CAM typically costs under €11, making it a highly scalable and cost efficient alternative for distributed deployments.
The system does not rely on a single climatological source.
Instead, it applies:
- Schema normalization (units, timestamp alignment → UTC)
- Temporal alignment (hour-level indexing)
- Bias inspection between grid-based and station data
- Missing value handling (interpolation or controlled forward-fill)
- Extreme-event preservation (no aggressive smoothing)
This produces structured multi-year historical datasets per city.
Potential causes:
- Model upgrades in Open-Meteo upstream sources.
- Climate trend shifts over decade-long horizon.
- Station equipment changes in Meteostat data.
Mitigation strategies:
- Rolling retraining
- Time-aware cross-validation
- Period-based performance benchmarking
- Drift monitoring (mean/std shift tracking)
All data queries are:
- Parameterized by latitude/longitude.
- Deterministic for fixed date ranges.
- Exported to versioned CSV artifacts.
- Stored under
data/anddata/processed/.
To reproduce datasets:
- Use documented API parameters.
- Fix date range.
- Maintain consistent variable list.
- Ensure UTC normalization before feature engineering.
The strategic goal is autonomy and robustness.
External climate APIs are used as authoritative historical references for:
- Learning climatological dynamics.
- Preserving realistic distributions.
- Modeling extreme weather behavior.
Internal ML models are trained on multi-source fused data to:
- Generate forecasts independently of external API uptime.
- Maintain resilience in case of third-party service failure.
- Improve long-term predictive fidelity through retraining.
The architecture is extensible: if additional official climate sources provide better signal quality, they can be integrated into the training pipeline.
Climate Loop leverages multiple authoritative climatological sources to build a resilient and autonomous forecasting system.
Two primary data providers are used:
| Source | Type | Role | Strength |
|---|---|---|---|
| Open-Meteo | Model/Reanalysis | Primary ML backbone | High temporal resolution |
| Meteostat | Station Observations | Validation & calibration | Real measured data |
Rather than relying on a single source, the project combines both to:
- Build multi-year historical datasets per city.
- Cross-validate modeled data against observed measurements.
- Detect systematic bias and distribution shifts.
- Preserve realistic extreme-event behavior.
Because no proprietary IoT infrastructure is currently deployed, these fused historical datasets are used to generate synthetic IoT-like hourly telemetry streams. This enables:
- Development and validation of ML forecasting models.
- Disaster severity classification training.
- Offline prediction capabilities independent of external APIs.
- A seamless transition path to real IoT data once available.
The architecture is designed to incorporate additional official climate datasets if they improve realism or predictive signal quality.
In essence:
External climate APIs provide authoritative historical signals.
Internal ML models are trained to learn and generalize those dynamics, enabling scalable, API-independent, production-ready forecasting.