A comprehensive, production-ready checklist for reviewing data validation, ETL testing, and automated reconciliation codebases.
Code reviews for data engineering tools need more rigor than standard web apps.
A subtle bug in a data validation framework can cause silent pipeline failures, false positive test passes, or accidental execution of unbounded SQL queries on production warehouses.
Whether you are building a custom data framework or maintaining automated ETL tests, use this generalized checklist during code reviews to keep your test suites secure, performant, and reliable.
1.
Test Case Configuration (YAML / JSON) TC ID Matching: Ensure the tc_id value matches the configuration filename exactly.
Schema Validity: Verify that type (e.g., count, data, recon, file) and source/target drivers are valid and supported.
Explicit Enablers: Confirm the enabled field is explicitly set (true or false) rather than omitted.
Relative File Paths: For file-based validation, ensure paths are relative to defined source/target data directories.
Non-Empty Queries: Confirm SQL sources and targets include non-empty query strings or valid template paths.
Unique Case IDs: Ensure test case identifiers are unique across the test suite directory.
Documented Rationale: If a test case has enabled: false or uses numeric tolerance thresholds (validation_tolerance), ensure a comment explains the business reason.
Dependency Order: Verify that basic structural checks (COUNT) run prior to deep comparisons (DATA / RECON).
2.
SQL & Query Logic Explicit Projections: No SELECT *.
All columns must be explicitly listed to avoid schema drift breaks.
Alignment: Source and target queries must return compatible data types and matching column ordering.
Environment Isolation: Check that query strings contain zero hardcoded hostnames, schema names, or environment paths.
Secret Hygiene: Ensure queries contain no hardcoded credentials or connection strings.
Warehouse Pushdown: Confirm filtering and heavy aggregation occur at the database level (WHERE / GROUP BY) rather than pulling full tables into application memory.
Determinism: Queries must return deterministic ordering (e.g., explicit ORDER BY on primary keys) so differential comparisons yield consistent results.
3.
Validator Core Logic Standardized Return Schema: Validator routines must always return a consistent payload schema (e.g., status, summary, src_row_count, tgt_row_count, matched_rows).
Status Consistency: Status values should use normalized uppercase strings ("PASS" / "FAIL").
Null/NaN Handling: Verify that NaN and NULL comparisons explicitly account for missing data parity.
Memory Guards: Row capping (e.g., .head(1000)) or chunking should be enforced to prevent OOM errors on large datasets.
Safe Numbers: Ensure tolerance thresholds enforce non-negative values (e.g., max(0.0, float(tolerance))).
Error Preservation: No silent exception swallowing — all except blocks must either re-raise or log through the logger framework.
4.
Execution & Pipeline Runners Type Dispatcher Safety: Ensure the execution runner raises an explicit ValueError when encountering an unsupported validation type.
Type Coercion Guards: Data loaders (e.g., CSV readers) should default string types (dtype=str) where appropriate to prevent silent type conversions (e.g., dropping leading zeroes in zip codes).
Isolation: Ensure individual test execution is wrapped in try/except blocks so a single failing test case doesn't collapse the entire run.
Output Capture: Redirect standard output streams properly so module logs are correctly captured in final HTML or JSON reports.
5.
Security & Data Safety No Hardcoded Credentials: Check that API keys, passwords, or connection strings are strictly retrieved from environment variables or secret managers.
No Raw PII: Confirm that local test fixtures (data/src, data/tgt) contain synthetic data instead of production PII or financial records.
Injection Prevention: SQL execution calls must use parameterized inp