Streaming read, convert, and inspect 100+ formats. Install iterabledata[mcp] and run iterable-mcp.
Iterable Data is a Python library for reading and writing data files row by row in a consistent, iterator-based interface. It provides a unified API for working with various data formats (CSV, JSON, Parquet, XML, etc.) similar to csv.DictReader but supporting many more formats.
This library simplifies data processing and conversion between formats while preserving complex nested data structures (unlike pandas DataFrames which require flattening).
fast, balanced, or max compression settings with effective-setting diagnosticswith statements for automatic resource cleanuptrust=True to acknowledge)fst extra; experimental).npy / .npz array rows (read & write; npy extra)zarr extra; experimental)atom_site loops; experimental).mat variables (mat extra; experimental)geophysical extra; experimental)geophysical extra; experimental)geophysical extra; experimental).geojsonl, .geojsons); streaming-friendlygeospatial extra; experimental)geospatial extra; experimental).asc)lidar extra; experimental)hdf5 extra; experimental)rdf extra; experimental).mdb / .accdb tables (access extra; experimental).123 spreadsheets (experimental)xml extra; experimental).vcf; requires bio extra)alignment extra and an explicit reference when needed)format="webdataset"; experimental)otlp extra)See the formats documentation (or docs/docs/formats/ in this repo) for per-format parameters, record shapes, and extras.
py7zr via iterabledata[compression])Python 3.10+
pip install iterabledata
The PyPI package is iterabledata. Import iterable:
from iterable import open_iterable
Or install from source:
git clone https://github.com/datenoio/iterabledata.git
cd iterabledata
pip install .
IterableData supports optional extras for additional features:
# AI-powered documentation generation
pip install iterabledata[ai]
# Database ingestion (PostgreSQL, ClickHouse, MongoDB, MySQL, Elasticsearch, etc.)
pip install iterabledata[db]
# RDF formats (TriG, N3, TriX)
pip install iterabledata[rdf]
# Excel Binary (XLSB)
pip install iterabledata[xlsb]
# Graph formats (GraphML, GEXF, DOT)
pip install iterabledata[graph]
# Alignment formats (BAM, SAM, CRAM)
pip install iterabledata[alignment]
# Genomic formats (VCF/BCF, CRAM, BED, GFF3/GTF via pysam and bio readers)
pip install iterabledata[bio]
# Zarr chunked array stores
pip install iterabledata[zarr]
# GeoParquet, FlatGeobuf, FileGDB, MapInfo MIF
pip install iterabledata[parquet,geospatial]
# LiDAR LAS point clouds
pip install iterabledata[lidar]
# MATLAB .mat files
pip install iterabledata[mat]
# SEG-Y, GRIB2, miniSEED
pip install iterabledata[geophysical]
# Microsoft Access (.mdb/.accdb)
pip install iterabledata[access]
# R fst frames (requires a suitable fst/rfst binding)
pip install iterabledata[fst]
# OpenTelemetry JSON and Protobuf export profiles
pip install iterabledata[otlp]
# Lakehouse table formats (Delta, Iceberg, Lance, Hudi, DuckLake)
pip install iterabledata[lakehouse]
# Apache Paimon (tables + Row + Mosaic file formats; separate from lakehouse)
pip install iterabledata[paimon]
# Or individually:
# pip install iterabledata[paimon-table]
# pip install iterabledata[paimon-row]
# pip install iterabledata[paimon-mosaic]
# pip install iterabledata[ducklake]
# Individual format extras (one per format family), for example:
pip install iterabledata[avro] # Apache Avro
pip install iterabledata[npy] # NumPy .npy/.npz
pip install iterabledata[ods] # OpenDocument spreadsheets
pip install iterabledata[rdata] # R RData/RDS
pip install iterabledata[ics] # iCalendar
# Also available: ubj, vcf, capnp, thrift, fbs, edn, hocon, der, bencode, ldif, hdf5, xml, rdf
# All optional dependencies
pip install iterabledata[all]
AI Features ([ai]): Enables AI-powered documentation generation using OpenAI, OpenRouter, Ollama, LMStudio, or Perplexity.
Database Engines ([db]): Enables read-only database access as iterable data sources. Supports PostgreSQL, ClickHouse, MySQL/MariaDB, Microsoft SQL Server, SQLite, MongoDB, and Elasticsearch/OpenSearch. Includes convenience groups:
[db-sql]: SQL databases only (PostgreSQL, ClickHouse, MySQL, MSSQL)[db-nosql]: NoSQL databases only (MongoDB, Elasticsearch)Genomic formats ([bio]): Enables genomic VCF/BCF, CRAM, BED, GFF3, and GTF support. Alignment formats use pysam and may require a reference file.
Geospatial / scientific extras: [geospatial] covers FileGDB and MapInfo MIF (plus existing GeoPackage/Shapefile stack). [lidar], [mat], and [geophysical] enable LAS, MATLAB MAT, and SEG-Y/GRIB2/miniSEED respectively. Many structure formats (XYZ, CIF, PDB, ASCII Grid, CZML, EDI, WebDataset, Lotus WK1) need no extra.
Lakehouse ([lakehouse]): Delta Lake, Apache Iceberg, Lance, Apache Hudi, and DuckLake. Delta, Iceberg, and DuckLake support bounded writes; Hudi is read-only for now.
Paimon ([paimon]): Apache Paimon warehouse tables plus Row and Mosaic file formats. Install [paimon-table], [paimon-row], or [paimon-mosaic] individually if you only need one surface. DuckLake alone is also available as [ducklake].
See the API documentation for details on these features.
For AI agents and LLM tooling, see llms.txt (short index), llms-full.txt (copy-paste recipes), the portable skill skills/iterabledata/SKILL.md, and CONTRIBUTING.md.
Generate dataset documentation with a local LLM (no API key) via LM Studio or Ollama, or use OpenAI:
from iterable.ai import doc
# Local (LM Studio on http://localhost:1234/v1)
documentation = doc.generate(
"data.csv",
provider="lmstudio",
base_url="http://localhost:1234/v1",
format="markdown",
)
# Or analyze structure + docs in one call
from iterable.ops import inspect
analysis = inspect.analyze("data.csv", autodoc=True, autodoc_provider="openai")
print(analysis["documentation"])
Need structured, machine-readable output? Use generate_blocks() to get independent documentation blocks (general, schema, quality, examples, statistics, agent_skill; plus opt-in codebook) plus the assembled markdown:
from iterable.ai import doc
result = doc.generate_blocks(
"data.csv",
provider="openai",
context={"title": "Sales 2025", "description": "Monthly sales export"},
progress=lambda event: print(event.stage, event.detail),
)
print(result["blocks"]["schema"]["data"]) # structured JSON
print(result["blocks"]["agent_skill"]["markdown"]) # portable agent skill (YAML + Markdown)
print(result["full_document_markdown"]) # assembled markdown
The agent_skill block emits a portable skill document (YAML frontmatter + Markdown) that AI agents can load for dataset-specific load/query/safety guidance.
Install AI support: pip install iterabledata[ai]. See examples/ai/ and docs/docs/api/ai.md.
from iterable.ops import schema, stats
sch = schema.infer("nested.jsonl", flatten_nested=True)
print(sch["fields"]["capital_city.lat"]["type"])
summary = stats.compute("nested.jsonl", flatten_nested=True)
print(summary["capital_city.lat"]["mean"])
For multi-table workbooks/databases, iterable.ai.table_profile.profile_selected_table()
profiles one named sheet/table under row/time budgets (nested flattening enabled).
from iterable.catalog import describe_format
info = describe_format("xml")
print(info["example_args"]) # {'tagname': 'item'}
Full export: dev/formats.json or export_catalog(format="json"). See docs/docs/api/catalog.md.
from iterable import open_iterable
with open_iterable("data.csv.gz") as source:
for row in source:
print(row)
from iterable import open_iterable
with open_iterable("output.jsonl.zst", mode="w") as dest:
for item in my_data:
dest.write(item)
from iterable import open_iterable
with open_iterable("data.csv.xz") as source:
n = 0
for row in source:
n += 1
if n % 1000 == 0:
print(f"Processed {n} rows")
from iterable import open_iterable
with open_iterable("data.jsonl") as source:
for row in source:
print(row)
with open_iterable("data.parquet") as source:
for row in source:
print(row)
with open_iterable("data.xml", iterableargs={"tagname": "item"}) as source:
for row in source:
print(row)
with open_iterable("data.xlsx") as source:
for row in source:
print(row)
# Read GeoJSON Text Sequence (streaming, one feature per line)
with open_iterable('features.geojsonl') as source:
for feature in source:
print(feature['properties'], feature['geometry'])
# Stream records from a TAR archive (members detected by filename)
with open_iterable('dataset.tar.gz', iterableargs={'members': '*.csv'}) as source:
for row in source:
print(row['_member'], row)
# Read genomic VCF (requires pip install iterabledata[bio])
with open_iterable('variants.vcf') as source:
for variant in source:
print(variant['chrom'], variant['pos'], variant['ref'], variant['alt'])
# Read genomic intervals (BED, GFF3, or GTF)
with open_iterable('genes.gff3') as source:
for feature in source:
print(feature['seqid'], feature['type'], feature['start'], feature['end'])
# Read a Zarr array (requires pip install iterabledata[zarr])
with open_iterable('signals.zarr', iterableargs={'array': 'values'}) as source:
for row in source:
print(row['value'])
# Read GeoParquet or FlatGeobuf (requires parquet/geospatial extras)
with open_iterable('roads.geoparquet') as source:
for feature in source:
print(feature.get('geometry'), feature.get('properties'))
# Read an OTLP JSON export (requires pip install iterabledata[otlp])
with open_iterable('telemetry.otlp.json') as source:
for item in source:
print(item['signal'], item['record'])
from iterable import open_iterable
# Read from PostgreSQL database
with open_iterable(
'postgresql://user:password@localhost:5432/mydb',
engine='postgres',
iterableargs={'query': 'users'}
) as source:
for row in source:
print(row)
# Read specific columns with filtering
with open_iterable(
'postgresql://localhost/mydb',
engine='postgres',
iterableargs={
'query': 'users',
'columns': ['id', 'name', 'email'],
'filter': 'active = TRUE'
}
) as source:
for row in source:
print(row)
# Read from ClickHouse database
with open_iterable(
'clickhouse://user:password@localhost:9000/analytics',
engine='clickhouse',
iterableargs={'query': 'events', 'settings': {'max_threads': 4}}
) as source:
for row in source:
print(row)
# Convert database to file
from iterable.convert import convert
convert(
fromfile='postgresql://localhost/mydb',
tofile='users.parquet',
iterableargs={'engine': 'postgres', 'query': 'users'}
)
# Convert ClickHouse to Parquet
convert(
fromfile='clickhouse://localhost:9000/analytics',
tofile='events.parquet',
iterableargs={'engine': 'clickhouse', 'query': 'events'}
)
from iterable import open_iterable
from iterable.helpers.detect import detect_file_type, detect_file_type_from_content
from iterable.helpers.utils import detect_encoding, detect_delimiter
# Detect file type and compression (uses filename extension)
result = detect_file_type('data.csv.gz')
print(f"Type: {result['datatype']}, Codec: {result['codec']}")
# Content-based detection (for files without extensions or streams)
with open('data.unknown', 'rb') as f:
detection_result = detect_file_type_from_content(f)
if detection_result:
format_id, confidence, method = detection_result
print(f"Detected format: {format_id} (confidence: {confidence:.2f}, method: {method})")
# open_iterable() automatically uses content-based detection as fallback
# Works with files without extensions, streams, or incorrect extensions
with open_iterable('data.unknown') as source: # Detects from content
for row in source:
print(row)
# Detect encoding for CSV files
encoding_info = detect_encoding('data.csv')
print(f"Encoding: {encoding_info['encoding']}, Confidence: {encoding_info['confidence']}")
# Detect delimiter for CSV files
delimiter = detect_delimiter('data.csv', encoding=encoding_info['encoding'])
# Open with detected settings
source = open_iterable('data.csv', iterableargs={
'encoding': encoding_info['encoding'],
'delimiter': delimiter
})
IterableData provides a comprehensive exception hierarchy and configurable error handling:
from iterable import open_iterable
from iterable.exceptions import (
FormatDetectionError,
FormatNotSupportedError,
FormatParseError,
ReadError,
CodecError,
IterableDataError,
)
# Basic exception handling
try:
with open_iterable('data.unknown') as source:
for row in source:
process(row)
except FormatDetectionError as e:
print(f"Could not detect format: {e.reason}")
# Try with explicit format or check file content
except FormatNotSupportedError as e:
print(f"Format '{e.format_id}' not supported: {e.reason}")
# Install missing dependencies or use different format
except FormatParseError as e:
print(f"Failed to parse {e.format_id} format")
if e.position:
print(f"Error at position: {e.position}")
except ReadError as e:
print(f"Read failed: {e}")
except IterableDataError as e:
print(f"Library error: {e}")
except CodecError as e:
print(f"Compression error with {e.codec_name}: {e.message}")
# Check file integrity or try different codec
except Exception as e:
print(f"Unexpected error: {e}")
Configurable Error Policies: Control how malformed records are handled:
# Skip malformed records and continue processing
with open_iterable(
'data.csv',
iterableargs={'on_error': 'skip', 'error_log': 'errors.log'}
) as src:
for row in src:
process(row) # Only processes valid rows
# Warn on errors but continue processing
with open_iterable(
'data.jsonl',
iterableargs={'on_error': 'warn', 'error_log': 'errors.log'}
) as src:
for row in src:
process(row) # Warnings logged, processing continues
# Default: raise exceptions immediately (existing behavior)
with open_iterable('data.csv', iterableargs={'on_error': 'raise'}) as src:
for row in src:
process(row)
No silent empty reads: Under the default policy (on_error='raise'), a malformed non-empty file raises FormatParseError rather than yielding zero records. Use on_error='skip' or 'warn' to tolerate bad records explicitly.
Pickle safety: Unpickling executes arbitrary code. Reading pickle files emits a warning unless you pass trust=True:
with open_iterable('data.pickle', iterableargs={'trust': True}) as source:
for row in source:
process(row)
Error Logging: Structured JSON logs with context (filename, row number, byte offset, error message, original line).
See Exception Hierarchy documentation for complete exception reference.
from iterable.helpers.capabilities import (
get_format_capabilities,
get_capability,
list_all_capabilities
)
# Get all capabilities for a format
caps = get_format_capabilities("csv")
print(f"CSV readable: {caps['readable']}")
print(f"CSV writable: {caps['writable']}")
print(f"CSV supports totals: {caps['totals']}")
print(f"CSV supports tables: {caps['tables']}")
# Query a specific capability
is_writable = get_capability("json", "writable")
has_totals = get_capability("parquet", "totals")
supports_tables = get_capability("xlsx", "tables")
# List capabilities for all formats
all_caps = list_all_capabilities()
for format_id, capabilities in all_caps.items():
if capabilities.get("tables"):
print(f"{format_id} supports multiple tables")
from iterable import open_iterable
from iterable.convert import convert
# Simple format conversion
convert('input.jsonl.gz', 'output.parquet')
# Convert with options
convert(
'input.csv.xz',
'output.jsonl.zst',
iterableargs={'delimiter': ';', 'encoding': 'utf-8'},
batch_size=10000
)
# Convert and flatten nested structures
convert(
'input.jsonl',
'output.csv',
is_flatten=True,
batch_size=50000
)
Use atomic writes to ensure output files are never left in a partially written state:
from iterable.convert import convert
from iterable.pipeline import pipeline
# Convert with atomic writes (production-safe)
result = convert('input.csv', 'output.parquet', atomic=True)
# Output file only appears when conversion completes successfully
# Atomic writes in pipelines
pipeline(
source=source,
destination=destination,
process_func=transform_func,
atomic=True # Ensures destination file is only created on success
)
Benefits: Prevents data corruption from crashes, interruptions, or mid-process failures. Original files are preserved on failure.
Convert multiple files at once using glob patterns, directories, or file lists:
from iterable.convert import bulk_convert
# Convert all CSV files matching glob pattern
result = bulk_convert('data/raw/*.csv.gz', 'data/processed/', to_ext='parquet')
# Convert with custom filename pattern
result = bulk_convert('data/*.csv', 'output/', pattern='{name}.parquet')
# Convert entire directory
result = bulk_convert('data/raw/', 'data/processed/', to_ext='parquet')
# Access results
print(f"Converted {result.successful_files}/{result.total_files} files")
print(f"Total rows: {result.total_rows_out}")
print(f"Throughput: {result.throughput:.0f} rows/second")
# Check individual file results
for file_result in result.file_results:
if file_result.success:
print(f"✓ {file_result.source_file}: {file_result.result.rows_out} rows")
else:
print(f"✗ {file_result.source_file}: {file_result.error}")
Features: Error resilience (continues if one file fails), aggregated metrics, flexible output naming with placeholders ({name}, {stem}, {ext}).
Track conversion and pipeline progress with callbacks, progress bars, and structured metrics:
from iterable.convert import convert
from iterable.pipeline import pipeline
# Progress callback for conversions
def progress_cb(stats):
print(f"Progress: {stats['rows_read']} rows read, "
f"{stats['rows_written']} rows written, "
f"{stats.get('elapsed', 0):.2f}s elapsed")
# Convert with progress tracking
result = convert(
'input.csv',
'output.parquet',
progress=progress_cb,
show_progress=True # Also shows tqdm progress bar
)
# Access conversion metrics
print(f"Converted {result.rows_out} rows in {result.elapsed_seconds:.2f}s")
print(f"Read {result.bytes_read} bytes, wrote {result.bytes_written} bytes")
# Pipeline with progress and metrics
result = pipeline(
source=source,
destination=destination,
process_func=transform_func,
progress=progress_cb # Progress callback
)
# Access pipeline metrics (supports both attribute and dict access)
print(f"Processed {result.rows_processed} rows")
print(f"Throughput: {result.throughput:.0f} rows/second")
print(f"Exceptions: {result.exceptions}")
# Backward compatible: result['rec_count'] also works
Features: Real-time progress callbacks, automatic progress bars with tqdm, structured metrics objects (ConversionResult, PipelineResult).
from iterable import open_iterable
from iterable.pipeline import pipeline
source = open_iterable('input.parquet')
destination = open_iterable('output.jsonl.xz', mode='w')
def transform_record(record, state):
"""Transform each record"""
# Add processing logic
out = {}
for key in ['name', 'email', 'age']:
if key in record:
out[key] = record[key]
return out
def progress_callback(stats, state):
"""Called every trigger_on records"""
print(f"Processed {stats['rec_count']} records, "
f"Duration: {stats.get('duration', 0):.2f}s")
def final_callback(stats, state):
"""Called when processing completes"""
print(f"Total records: {stats['rec_count']}")
print(f"Total time: {stats['duration']:.2f}s")
result = pipeline(
source=source,
destination=destination,
process_func=transform_record,
trigger_func=progress_callback,
trigger_on=1000,
final_func=final_callback,
start_state={},
atomic=True # Use atomic writes for production safety
)
# Access pipeline metrics
print(f"Throughput: {result.throughput:.0f} rows/second")
source.close()
destination.close()
from iterable.datatypes.jsonl import JSONLinesIterable
from iterable.datatypes.bsonf import BSONIterable
from iterable.codecs.gzipcodec import GZIPCodec
from iterable.codecs.lzmacodec import LZMACodec
# Read gzipped JSONL
read_codec = GZIPCodec('input.jsonl.gz', mode='r', open_it=True)
reader = JSONLinesIterable(codec=read_codec)
# Write LZMA compressed BSON
write_codec = LZMACodec('output.bson.xz', mode='wb', open_it=False)
writer = BSONIterable(codec=write_codec, mode='w')
for row in reader:
writer.write(row)
reader.close()
writer.close()
Read and write data directly from cloud object storage (S3, GCS, Azure):
from iterable import open_iterable
# Read from S3
with open_iterable('s3://my-bucket/data/events.csv') as source:
for row in source:
print(row)
# Read compressed file from GCS
with open_iterable('gs://my-bucket/data/events.jsonl.gz') as source:
for row in source:
process(row)
# Write to Azure Blob Storage
with open_iterable(
'az://my-container/output/results.jsonl',
mode='w',
iterableargs={'storage_options': {'connection_string': '...'}}
) as dest:
dest.write({'name': 'Alice', 'age': 30})
dest.write({'name': 'Bob', 'age': 25})
Supported Providers:
s3:// and s3a:// schemesgs:// and gcs:// schemesaz://, abfs://, and abfss:// schemesInstallation: pip install iterabledata[cloud]
Note: DuckDB engine does not support cloud storage URIs; use engine='internal' (default).
The DuckDB engine provides high-performance querying with advanced optimizations:
from iterable import open_iterable
# Basic DuckDB usage
source = open_iterable('data.csv.gz', engine='duckdb')
total = source.totals() # Fast counting
for row in source:
print(row)
source.close()
# Column projection pushdown (only read specified columns)
with open_iterable(
'data.csv',
engine='duckdb',
iterableargs={'columns': ['name', 'age']} # Reduces I/O and memory
) as src:
for row in src:
process(row)
# Filter pushdown (filter at database level)
with open_iterable(
'data.csv',
engine='duckdb',
iterableargs={'filter': "age > 18 AND status = 'active'"}
) as src:
for row in src:
process(row)
# Combined column projection and filtering
with open_iterable(
'data.parquet',
engine='duckdb',
iterableargs={
'columns': ['name', 'age', 'email'],
'filter': 'age > 18'
}
) as src:
for row in src:
process(row)
# Direct SQL query support
with open_iterable(
'data.parquet',
engine='duckdb',
iterableargs={
'query': 'SELECT name, age FROM read_parquet(\'data.parquet\') WHERE age > 18 ORDER BY age DESC LIMIT 100'
}
) as src:
for row in src:
process(row)
Supported Formats: CSV, JSONL, NDJSON, JSON, Parquet
Supported Codecs: GZIP, ZStandard (.zst)
Benefits: Reduced I/O, lower memory usage, faster processing through database-level optimizations
from iterable import open_iterable
source = open_iterable('input.jsonl')
destination = open_iterable('output.parquet', mode='w')
# Read and write in batches for better performance
batch = []
for row in source:
batch.append(row)
if len(batch) >= 10000:
destination.write_bulk(batch)
batch = []
# Write remaining records
if batch:
destination.write_bulk(batch)
source.close()
destination.close()
Use native batches when both endpoints are Parquet or Arrow/Feather. The selection is pushed into the columnar reader when supported; otherwise the conversion safely falls back to the regular row/bulk loop:
from iterable.convert import BatchSelection, convert
convert(
'events.parquet',
'events-copy.parquet',
use_native_batch=True,
selection=BatchSelection(columns=('event_id', 'created_at'), batch_size=8192),
toiterableargs={'row_group_size': 32768},
)
For general compression workloads, use the balanced profile. Choose fast
for CPU-bound pipelines or max for archival output:
with open_iterable('events.jsonl.zst', mode='w', codecargs={'profile': 'fast'}) as dest:
dest.write_bulk(records)
See native batches, codec profiles, and the performance guide for selection, memory, row-group, and fallback behavior.
from iterable import open_iterable
from iterable.ai.fileinfo import open_table
# Read Excel file (specify sheet or page)
xls_file = open_iterable('data.xlsx', iterableargs={'page': 0})
for row in xls_file:
print(row)
xls_file.close()
# Open a named sheet (uses page index under the hood)
sheet = open_table('data.xlsx', 'Sheet2')
for row in sheet:
print(row)
sheet.close()
from iterable import open_iterable
# Parse XML with specific tag name
xml_file = open_iterable(
'data.xml',
iterableargs={
'tagname': 'book',
'prefix_strip': True # Strip XML namespace prefixes
}
)
for item in xml_file:
print(item)
xml_file.close()
Convert iterable data to Pandas, Polars, or Dask DataFrames:
from iterable import open_iterable
# Convert to Pandas DataFrame
with open_iterable('data.csv.gz') as source:
df = source.to_pandas()
print(df.head())
# Chunked processing for large files
with open_iterable('large_data.csv') as source:
for df_chunk in source.to_pandas(chunksize=100_000):
# Process each chunk
result = df_chunk.groupby('category').sum()
process_chunk(result)
# Convert to Polars DataFrame
with open_iterable('data.csv.gz') as source:
df = source.to_polars()
print(df.head())
# Convert to Dask DataFrame (single file)
with open_iterable('data.csv.gz') as source:
ddf = source.to_dask()
result = ddf.groupby('category').sum().compute()
# Multi-file Dask DataFrame (automatic format detection)
from iterable.helpers.bridges import to_dask
ddf = to_dask(['file1.csv', 'file2.jsonl', 'file3.parquet'])
result = ddf.groupby('category').sum().compute()
Note: DataFrame bridges require optional dependencies. Install with:
pip install iterabledata[dataframes] # All DataFrame libraries
# Or individually:
pip install pandas
pip install polars
pip install "dask[dataframe]"
IterableData includes complete type annotations and typed helper functions for modern Python development:
from iterable import open_iterable
from iterable.helpers.typed import as_dataclasses, as_pydantic
from dataclasses import dataclass
from pydantic import BaseModel
# Type aliases for better code readability
from iterable.types import Row, IterableArgs, CodecArgs
# Convert to dataclasses for type safety
@dataclass
class Person:
name: str
age: int
email: str | None = None
with open_iterable('people.csv') as source:
for person in as_dataclasses(source, Person):
# Full IDE autocomplete and type checking
print(person.name, person.age)
# Convert to Pydantic models with validation
class PersonModel(BaseModel):
name: str
age: int
email: str | None = None
with open_iterable('people.jsonl') as source:
for person in as_pydantic(source, PersonModel, validate=True):
# Automatic schema validation
print(person.name, person.age)
# Access as Pydantic model with all features
Benefits:
py.typed marker file enables mypy, pyright, and other type checkersInstallation: pip install iterabledata[pydantic] for Pydantic support
from iterable.datatypes.xml import XMLIterable
from iterable.datatypes.parquet import ParquetIterable
from iterable.codecs.bz2codec import BZIP2Codec
# Read compressed XML
read_codec = BZIP2Codec('data.xml.bz2', mode='r')
reader = XMLIterable(codec=read_codec, tagname='page')
# Write to Parquet with schema adaptation
writer = ParquetIterable(
'output.parquet',
mode='w',
use_pandas=False,
adapt_schema=True,
batch_size=10000
)
batch = []
for row in reader:
batch.append(row)
if len(batch) >= 10000:
writer.write_bulk(batch)
batch = []
if batch:
writer.write_bulk(batch)
reader.close()
writer.close()
open_iterable(filename, mode='r', engine='internal', codecargs={}, iterableargs={})Opens a file and returns an iterable object.
Parameters:
filename (str): Path to the file (supports local files and cloud storage URIs: s3://, gs://, az://)mode (str): File mode ('r' for read, 'w' for write)engine (str): Processing engine ('internal' or 'duckdb')codecargs (dict): Arguments for codec initializationiterableargs (dict): Arguments for iterable initialization
columns (list[str]): For DuckDB engine, only read specified columns (pushdown optimization)filter (str | callable): For DuckDB engine, filter rows at database level (SQL string or Python callable)query (str): For DuckDB engine, execute custom SQL query (read-only)on_error (str): Error policy ('raise', 'skip', or 'warn')error_log (str | file-like): Path or file object for structured error loggingstorage_options (dict): Cloud storage authentication optionsReturns: Iterable object for the detected file type
detect_file_type(filename)Detects file type and compression codec from filename.
Returns: Dictionary with success, datatype, and codec keys
convert(fromfile, tofile, iterableargs={}, toiterableargs={}, scan_limit=1000, batch_size=50000, silent=True, is_flatten=False, use_totals=False, progress=None, show_progress=False, atomic=False, use_native_batch=False, selection=None, strict_native=False) -> ConversionResultConverts data between formats.
Parameters:
fromfile (str): Source file pathtofile (str): Destination file pathiterableargs (dict): Options for reading source filetoiterableargs (dict): Options for writing destination filescan_limit (int): Number of records to scan for schema detectionbatch_size (int): Batch size for bulk operationssilent (bool): Suppress progress outputis_flatten (bool): Flatten nested structuresuse_totals (bool): Use total count for progress tracking (if available)progress (callable): Optional callback function receiving progress stats dictionaryshow_progress (bool): Display progress bar using tqdm (if available)atomic (bool): Write to temporary file and atomically rename on successuse_native_batch (bool): Request native columnar batch transfer when both endpoints support itselection (BatchSelection or dict): Optional columns, row range, slice, table, or backend predicate selectionstrict_native (bool): Raise instead of falling back when native batching or the requested selection is unsupportedReturns: ConversionResult object with:
rows_in (int): Total rows readrows_out (int): Total rows writtenelapsed_seconds (float): Conversion timebytes_read (int | None): Bytes read (if available)bytes_written (int | None): Bytes written (if available)errors (list[Exception]): List of errors encounteredbulk_convert(source, destination, pattern=None, to_ext=None, **kwargs) -> BulkConversionResultConvert multiple files at once using glob patterns, directories, or file lists.
Parameters:
source (str): Glob pattern, directory path, or file pathdestination (str): Output directory or filename patternpattern (str): Filename pattern with placeholders ({name}, {stem}, {ext})to_ext (str): Replace file extension (e.g., 'parquet')**kwargs: All parameters from convert() functionReturns: BulkConversionResult object with:
total_files (int): Total files processedsuccessful_files (int): Files successfully convertedfailed_files (int): Files that failedtotal_rows_in (int): Total rows read across all filestotal_rows_out (int): Total rows written across all filestotal_elapsed_seconds (float): Total conversion timefile_results (list[FileConversionResult]): Per-file resultserrors (list[Exception]): All errors encounteredthroughput (float | None): Rows per secondpipeline(source, destination, process_func, trigger_func=None, trigger_on=1000, final_func=None, reset_iterables=True, skip_nulls=True, start_state=None, debug=False, batch_size=1000, progress=None, atomic=False) -> PipelineResultExecute a data processing pipeline.
Parameters:
source (BaseIterable): Source iterable to read fromdestination (BaseIterable | None): Destination iterable to write toprocess_func (callable): Function to process each recordtrigger_func (callable | None): Function called periodically during processingtrigger_on (int): Number of records between trigger function callsfinal_func (callable | None): Function called after processing completesreset_iterables (bool): Reset iterables before processingskip_nulls (bool): Skip None results from process_funcstart_state (dict | None): Initial state dictionarydebug (bool): Raise exceptions instead of catching thembatch_size (int): Number of records to batch before writingprogress (callable | None): Optional callback function for progress updatesatomic (bool): Use atomic writes if destination is a fileReturns: PipelineResult object with:
rows_processed (int): Total rows processedelapsed_seconds (float): Processing timethroughput (float | None): Rows per secondexceptions (int): Number of exceptions encounterednulls (int): Number of null resultsresult.rows_processed) and dictionary access (result['rec_count']) for backward compatibilityAll iterable objects support:
read(skip_empty=True) -> Row - Read single recordread_bulk(num=DEFAULT_BULK_NUMBER) -> list[Row] - Read multiple recordswrite(record) - Write single recordwrite_bulk(records) - Write multiple recordsreset() - Reset iterator to beginningclose() - Close file handlesto_pandas(chunksize=None) - Convert to pandas DataFrame (optional chunked processing)to_polars(chunksize=None) - Convert to Polars DataFrame (optional chunked processing)to_dask(chunksize=1000000) - Convert to Dask DataFramelist_tables(filename=None) -> list[str] | None - List available tables/sheets/datasetshas_tables() -> bool - Check if format supports multiple tablesas_dataclasses(iterable, dataclass_type, skip_empty=True) -> Iterator[T]Convert dict-based rows from an iterable into dataclass instances.
Parameters:
iterable (BaseIterable): The iterable to read rows fromdataclass_type (type[T]): The dataclass type to convert rows toskip_empty (bool): Whether to skip empty rowsReturns: Iterator of dataclass instances
as_pydantic(iterable, model_type, skip_empty=True, validate=True) -> Iterator[T]Convert dict-based rows from an iterable into Pydantic model instances.
Parameters:
iterable (BaseIterable): The iterable to read rows frommodel_type (type[T]): The Pydantic model type to convert rows toskip_empty (bool): Whether to skip empty rowsvalidate (bool): Whether to validate rows against the model schemaReturns: Iterator of Pydantic model instances
Raises: ImportError if pydantic is not installed
to_dask(files, chunksize=1000000, **iterableargs) -> DaskDataFrameConvert multiple files to a unified Dask DataFrame with automatic format detection.
Parameters:
files (str | list[str]): Single file path or list of file pathschunksize (int): Number of rows per partition**iterableargs: Additional arguments to pass to open_iterable() for each fileReturns: Dask DataFrame containing data from all files
Raises: ImportError if dask or pandas is not installed
The internal engine uses pure Python implementations for all formats. It supports all file types and compression codecs.
The DuckDB engine provides high-performance querying capabilities for supported formats:
Use engine='duckdb' when opening files:
source = open_iterable('data.csv.gz', engine='duckdb')
See the examples directory for more complete examples:
simplewiki/ - Processing Wikipedia XML dumpsSee the tests directory for comprehensive usage examples and test cases.
Contributors: run the full suite with pytest --verbose. The performance regression gate is opt-in:
pytest tests/test_performance_regression.py -m performance --no-cov
Regenerate baselines intentionally after legitimate performance changes:
pytest tests/test_performance_regression.py -m performance --no-cov --update-baselines
See AGENTS.md for development conventions.
IterableData can be integrated with AI platforms and frameworks for intelligent data processing:
Source-derived launch command. Check the maintainer’s required arguments and credentials before running:
uvx iterabledataMerge this template into ~/Library/Application Support/Claude/claude_desktop_config.json. Keep existing servers. Add any arguments, credentials, and permissions required by the maintainer; this template has not been install-tested.
{
"mcpServers": {
"io-github-datenoio-iterabledata": {
"command": "uvx",
"args": [
"iterabledata"
]
}
}
}Restart Claude Desktop completely for changes to take effect. Confirm the server appears connected in the client’s tool list, then try a read-only example from its documentation.
Claude Desktop setup referenceIterableData works with any MCP-compatible client. Copy the config snippet from the Configuration section above and add it to the file shown for your client, then restart the application.
~/Library/Application Support/Claude/claude_desktop_config.jsonRestart Claude Desktop completely for changes to take effect.~/.cursor/mcp.jsonRestart Cursor for changes to take effect..vscode/mcp.jsonReload VS Code window for changes to take effect.~/.codeium/windsurf/mcp_config.jsonRestart Windsurf for changes to take effect..mcp.jsonSave at the project root, then start Claude Code in that project and review the MCP server approval prompt. Keep real credentials out of shared files.