Swarm CLI: aiswarm (on PATH; make install-aiswarm from the nudge repo).
Read workflow first:
aiswarm— common commands cheat sheetaiswarm instructions overview— required agent briefingaiswarm instructions handoff/tasks— peer send and backlog dispatchaiswarm this— this swarm's config + runtime.json path
After start, machine map (not git): /tmp/nudge-swarm/xml-iterator/runtime.json
Config: .aiswarm/config.yaml (cwd walk-up), $AISWARM_CONFIG, or explicit path.
Messaging: aiswarm send <pane> "msg" (durable log). Do NOT raw tmux send-keys.
Fast XML parser with streaming iterator interface, built in Rust with Python bindings.
Primary goal: defeat the infinite depth attack (content under outer elements that stay open
until late/EOF — FIRDS-like). See backlog/docs/streaming-memory-model-and-landscape.md.
xml_iterator/
├── src/lib.rs # Rust core: XMLIterator + Python bindings
├── xml_iterator/core.py # Python utilities: xml_to_dict, get_edge_counts
├── tests/ # Comprehensive pytest suite
└── benchmark*.py # Performance testing vs xmltodict
- XMLIterator: Streaming XML parser using quick-xml
- Events:
start,end,text,empty(self-closing tags),attr(opt-in) - Python bindings: PyO3 integration
- Protection: No depth limits - user controls via early termination
- Errors: malformed XML and undecodable text raise
ValueError, not silent truncation
iter_xml(path, attributes=False): Stream events(count, event, value); withattributes=Truealso yields('attr', (name, value))eventsxml_to_dict(path): Full-document dict (xmltodict-compatible). Modest files / parity only — not the multi-GB FIRDS path (rebuilds the tree → loses to infinite depth at that scale)get_edge_counts(path): Count tag hierarchies
✅ Streaming under open ancestors - child end while wrappers stay open (FIRDS shape)
✅ Matches xmltodict output - including attributes (namespace prefixes are stripped; see limitations)
✅ Early-termination streaming - stopping early avoids a full-document parse, a property of any streaming parser, not unique to this library
✅ Bounded memory if work is discarded - stream alone is not enough if every child is kept under open parents
✅ Real-world tested - handles 300MB+ ESMA FIRDS XML files via streaming
✅ Fails loudly - malformed XML and undecodable text raise ValueError (no silent truncation)
IMPORTANT: always benchmark a release build (make develop); debug builds are ~9x slower
and historically poisoned this project's numbers.
Release build, 2026-07-17 (see PERF_2026-07-17.md for the full story):
| Scenario | xml_iterator | baseline | Ratio |
|---|---|---|---|
| xml_to_dict (Rust-built as of TASK-2), synthetic 5000 items (1.8 MB) | 0.027s | xmltodict 0.157s | 5.8x faster |
| xml_to_dict, SwissProt 110 MB (results identical) | 3.0s | xmltodict 13.2s | 4.5x faster |
| stream drain synthetic 2k books (~50k events) | 0.013s | et 0.043s / sax 0.044s / lxml 0.032s | ~2.5–3.5x faster |
| iter_xml full drain, SwissProt (8.0M events) | 2.6s | ET.iterparse 6.6s | 2.5x faster |
| Rust get_edge_counts, SwissProt | 1.3s | - | aggregation stays in Rust |
| Early termination (stop at 1000 events) | 0.001s | N/A | early exit avoids full parse - any streaming parser gets this |
Stream backends: xml_iterator.comparators (et_iterparse, sax, lxml_iterparse).
Shared helpers in bench_common.py (measure once → print → JSON → README md).
make benchmark / make benchmark-real always include every stream backend or record a skip.
Before-today (v0.1.4) baseline for the same scenarios: xml_to_dict 1.1x vs xmltodict and SwissProt 11.6s (attribute-less, debug builds); iter_xml drain 20.5s (debug). See PERF_2026-07-17.md.
# Build and install
make develop
# Run tests
make test # All tests
make test-fast # Skip slow tests
# Run benchmarks
make benchmark # Synthetic data vs xmltodict
make benchmark-real # Real-world SwissProt XML data
make benchmark-firds # Real ESMA FIRDS data (downloads 17MB)
make benchmark-all # Run both real-world benchmarks
# Test specific components
pytest tests/test_basic.py # Core functionality
pytest tests/test_xmltodict.py # Compatibility
pytest tests/test_performance.py # Regression tests- Core functionality: streaming iterator, xmltodict-matching dict conversion, edge counting
- Tested: synthetic data, real-world XML files, adversarial/edge cases
- Benchmarked: performance measured vs xmltodict and stdlib
ET.iterparse
src/lib.rs: Main Rust implementationxml_iterator/core.py: Python utilities and xml_to_dicttests/test_xmltodict.py: Compatibility verificationbenchmark_real_world.py: Real-world performance testingbenchmark.py: Synthetic benchmarks
- Namespace prefixes stripped: only local tag/attribute names are kept
- Single file input: No streaming from network/pipes (file paths only)
- Python-only bindings: No other language bindings yet
Definition (project sense): useful content sits under outer elements that do not close until much later (often EOF). Tree/DOM must wait for outer close or retain the open tree. FIRDS dumps are the concrete instance: wrappers open near the start, millions of records at that depth, outers close at EOF.
<root> <Payload> <RefData> ← open almost whole file
<FinInstrm>...</FinInstrm> ← record complete; outers still open
...
</RefData></Payload></root> ← close near EOF
Protection: stream events at depth under open outers; process on child end; user discards
finished work and may early-stop. Not primarily max_depth caps (those are nesting/work
limits). Full xml_to_dict is modest files / parity only.
Regression: tests/test_firds_shape.py. Example: examples/firds_shape_stream.py.
- Rust: quick-xml, pyo3, encoding_rs_io
- Python: Standard library only (tests require pytest, xmltodict)
- Build: maturin for Python extension compilation
- Exact compatibility: matches xmltodict output including attributes
- Real-world data: ESMA FIRDS regulatory XML files
- Performance regression: Ensure no slowdowns
- Fail loudly: malformed XML and undecodable text raise
ValueError, no silent truncation
<CRITICAL_INSTRUCTION>
This project uses Backlog.md for task and project management.
For every user request in this project, run backlog instructions overview before answering or taking action.
Use the overview to decide whether to search, read, create, or update Backlog tasks.
Before task lifecycle actions, read the matching detailed guide:
backlog instructions task-creationbefore creating or splitting tasksbacklog instructions task-executionbefore planning, changing status or assignee, adding a plan or implementation notes, or implementing task workbacklog instructions task-finalizationbefore checking acceptance criteria, writing final summaries, or moving tasks to terminal statuses
Use backlog <command> --help before running unfamiliar commands. Help shows options, fields, and examples.
Do not edit Backlog task, draft, document, decision, or milestone markdown files directly. Use the backlog CLI so metadata, relationships, and history stay consistent.
</CRITICAL_INSTRUCTION>