Loading Data¶
CGMPy provides several ways to ingest CGM data, from a one-line helper to fully controlled per-device loaders. See Supported devices for the list of auto-detected formats.
One-line loader¶
The GlucoseAnalysis facade takes a path and does everything:
Internally it:
- Detects the file format (CSV / Parquet / DataFrame).
- Picks the right
DataLoader(generic or device-specific). - Processes and validates the data (numeric coercion, dedup, gap detection).
- Computes every metric.
- Makes the data available for plotting.
Modular loaders¶
If you want to control the pipeline step by step, use the modular classes.
GlucoseData¶
from cgmpy import GlucoseData
# From CSV
data = GlucoseData("data.csv")
# From Parquet (fast)
data = GlucoseData("data.parquet")
# From a DataFrame
import pandas as pd
data = GlucoseData(pd.read_csv("data.csv"))
Glucose units (mg/dL vs mmol/L)¶
CGMPy computes every metric in mg/dL. If your data is in mmol/L, pass
unit="mmol/L" and it is converted to mg/dL on load, so all metrics are
correct:
# mmol/L data (common outside the US) is normalized to mg/dL internally
data = GlucoseData("data_mmol.csv", unit="mmol/L")
data.unit # GlucoseUnit.MG_DL (canonical internal unit)
data.source_unit # GlucoseUnit.MMOLL (what you provided)
data.glucose_in_unit("mmol/L") # view the values back in mmol/L
DataLoader (low-level)¶
For maximum control, use DataLoader directly:
from cgmpy.data import DataLoader, DataProcessor, DataAnalyzer
loader = DataLoader()
processor = DataProcessor()
analyzer = DataAnalyzer()
raw = loader.load_from_csv("data.csv", time_col="time", glucose_col="glucose")
processed, diffs = processor.process_data(raw, "time", "glucose")
interval = analyzer.calculate_typical_interval(diffs)
print(f"Typical interval: {interval} min")
Device-specific loaders¶
| Device | Class | Notes |
|---|---|---|
| Dexcom Clarity | Dexcom |
Auto-detected from header. |
| FreeStyle Libre | Libreview |
Pass header=2 to skip metadata. |
| Medtronic CareLink | MedtronicCarelink |
|
| Tandem t:slim | TandemDiabetes |
Auto-detection:
from cgmpy.data import detect_device_type, create_specialized_loader
device = detect_device_type("data.csv") # 'dexcom', 'libreview', ...
loader = create_specialized_loader("data.csv")
Filtering¶
data = GlucoseData("data.csv")
# By date range
jan = data.filter_by_date_range("2024-01-01", "2024-01-31")
# By glucose range (e.g., physiological sanity)
normal = data.filter_by_glucose_range(40, 400)
Inspection¶
print(data) # __str__ shows a summary
info = data.info(include_disconnections=True)
print(info["n_records"], info["typical_interval"])
quality = data.get_data_quality_metrics()
print(quality["total_gaps"], quality["max_gap_hours"])
Exporting¶
data.to_parquet("optimized.parquet") # recommended for big data
data.to_csv("cleaned.csv")
data.to_excel("report.xlsx")
The Parquet writer preserves dtypes, uses snappy compression, and is ~10× faster to re-read than CSV.
Troubleshooting¶
All CGMPy-raised errors inherit from cgmpy.errors.CGMPyError. The
subclasses below are the ones you will encounter most often in the loading
path.
ColumnNotFoundError¶
Your CSV is missing the time column (or whichever column CGMPy expected for
the auto-detected device). Either it is not a recognized device export, or
you need to pass date_col=... and glucose_col=... to GlucoseData.
from cgmpy import CGMPyError, ColumnNotFoundError, GlucoseData
try:
data = GlucoseData("export.csv")
except ColumnNotFoundError as e:
print(f"Missing: {e.column}; available: {e.available}")
data = GlucoseData(
"export.csv",
date_col="Datetime",
glucose_col="Sensor Glucose (mg/dL)",
)
InvalidCSVFormatError¶
The CSV could not be parsed. Try specifying delimiter=';' (some locales
export with semicolons), or check that the file is not corrupted / not
actually a text file saved with the .csv extension.
from cgmpy import InvalidCSVFormatError, GlucoseData
try:
data = GlucoseData("export.csv", delimiter=";")
except InvalidCSVFormatError as e:
print(f"Cannot parse {e.file_path}: {e.reason}")
DeviceDetectionError¶
We could not auto-detect the device. The exception message lists the
columns that were found. Pass the loader class explicitly
(e.g. Dexcom("file.csv")), or use GlucoseData with
date_col= / glucose_col=.
from cgmpy import Dexcom, DeviceDetectionError, GlucoseData
try:
data = GlucoseData("export.csv")
except DeviceDetectionError as e:
print(f"Found columns: {e.columns_found}")
data = Dexcom("export.csv") # try Dexcom explicitly
EmptyDataError¶
No rows remained after filtering. Check that your start_date /
end_date range actually overlaps with the file, and that the file is
not itself empty.
from cgmpy import EmptyDataError
try:
jan = data.filter_by_date_range("2024-01-01", "2024-01-31")
except EmptyDataError as e:
print(f"Empty after: {e.context}")
AgataNotInstalledError¶
You need py_agata for the AGATA comparison. Install the optional
[agata] extra:
from cgmpy import AgataNotInstalledError
try:
analysis.run_agata_comparison()
except AgataNotInstalledError:
print("Run: pip install 'cgmpy[agata]'")
For the full error hierarchy, see cgmpy.errors in the
API reference.
See also¶
- Supported devices — auto-detected formats.
- Data formats — required columns and aliases.
- API reference → Data — every class and method.