Colby's Data Movers

Pathfinder

An ETL platform in Rust, being built to be open source.

The same lane as Meltano and dbt: local-first, transform-heavy, run on your own hardware. What I have not found in either is an audit trail that arrives already queryable. Findings land in DuckDB, so “prove this migration was clean” is a SQL query against the run rather than a support ticket. The model is open source plus consulting. There is no SaaS to buy and no vendor, us included, in your chain of custody.

Discuss your migration How it compares
8,931 tests, all passing $0 data egress
Tyler Colby Salesforce Certified Data Architecture & Management Designer
20+ years in Salesforce Migrations for St. Jude, NRDC & Houston Food Bank 10+ TB for Fortune 100 & regulated orgs

Three surfaces, one engine

The shape dbt settled on, and the right one: the engine does the work, and you reach it from whichever surface suits the job. Two of the three ship today. The third is the one I am building toward, and I would rather say which is which than let a diagram imply otherwise.

Studio · ships today

Author the mapping by hand.

The desktop workbench. Ranked field-mapping suggestions you accept or reject, live data preview, undo and redo, provenance on every transform, and a problems dock that tells you what will fail before you run it.

MCP · ships today

Let an agent drive it.

An MCP server exposing 73 tools across 18 domains, so Claude Code or Cursor can open a workspace, discover a schema, propose a mapping, typecheck the transforms and read back findings without a human in the UI. I have not found another ETL tool that treats agent operation as a first-class surface.

CLI · in progress

Put it in a pipeline.

A single pathfinder binary is the designed third surface and is not built yet. What exists today is pf-audit, which runs audits and exports JSON, Markdown or HTML from a terminal. When the unified binary lands it will be announced here, not implied.

SalesforceDynamics 365HubSpotCSVExcelJSONXMLYAML-defined SalesforceDynamics 365HubSpotCSVExcelJSONXMLYAML-defined
Pathfinder platform showing Fabric AI Hub with a large pattern library, DAG pipeline studio, audit checks, and Salesforce connectors

Two engines

Migration Engine

Move data between systems.

8 connectors. 97 transform functions. 8-stage pipeline with adaptive batching, checkpoint/resume, conflict detection, self-healing remediation, and five quality gates. Salesforce Bulk API v2 for high-volume loads.

Audit EngineNew

Know your data before you move it.

DuckDB-powered analytical engine running 45 checks across security, compliance, data quality, and architecture. Cloud detection for 14 cloud types. Vertical-specific checks. JSON, Markdown, and HTML report export. CLI or desktop.

45
Audit checks
8
Connectors
AI
assisted field mapping
8,931
Tests, all passing

New: Audit Engine

45 checks. DuckDB under the hood. Results in minutes.

Connect to any Salesforce org and run a comprehensive audit. 45 checks across five domains cover data quality, compliance, security, schema and metadata, and performance. Vertical rules for industries like nonprofits sit behind the audit-store build. The DuckDB analytical engine processes org metadata at columnar speed, not row-by-row. Cloud detection identifies 14 cloud types (Sales, Service, Nonprofit, Health, Financial Services, and more) and auto-selects relevant checks.

pf-audit runs audits from a terminal and exports JSON, Markdown or HTML. Findings land in DuckDB, so you can query them in SQL.

Pathfinder showing connected Salesforce org with audit-ready status and multiple data sources

Five gates between your data and the target.

Every record passes five checkpoints. A failed gate stops the migration rather than pausing it, so you fix the issue, re-run, and every record passes clean.

Pass
Pre-Extract
→
Pass
Post-Extract
→
Pass
Post-Transform
→
Pass
Pre-Load
→
Pass
Post-Load

Field mapping that takes an afternoon

AutoMapper scores every possible source-to-target field mapping using name similarity, type compatibility, and domain-specific patterns. Accept the suggestions, adjust with 97 built-in transform functions, or load a pre-built accelerator. Lookup tables, aggregations (SUM, COUNT, AVG with GROUP BY), and row context variables ($ROW_NUMBER, $BATCH_NUMBER) handle the edge cases.

97 transforms. Lookup tables. Undo/redo. Live data preview.

Mapping studio showing object mappings with field coverage percentages, quality scores, and transform expressions

Migrations that fix their own mistakes.

A rejected record is a problem to solve, not a line in a log. Missing required field, so it applies the default and retries. String over the field length, so it truncates. Rate limited, so it switches to Bulk API. Lookup target does not exist yet, so the record goes in a queue and gets a second pass once it does. Nine rules, and between them they cover the four failures that end most migrations at 2am.

9 remediation actions. Checkpoint/resume. Adaptive batch sizing (50 to 2,000 records).

Migration dashboard showing completed migrations with record counts, running migration with real-time progress, and error remediation

Field mapping the tool suggests and you approve

AutoMapper ranks every candidate field pairing on name similarity, type compatibility, and known Salesforce field shapes. The ranking is heuristic rather than a model call, so it is deterministic and it runs offline. You accept, adjust, or reject each one. Where a model is involved, it runs against Claude or a local Ollama model, and nothing leaves your machine unless you point it at a hosted provider.

Deterministic ranking. Offline by default. Every suggestion needs a human yes.

Pathfinder mapping studio showing ranked field-mapping suggestions with accept and adjust controls

The four ways migrations go wrong

"We exported 50,000 records and ended up with 12,000 duplicates."

Pathfinder's identity resolution uses Jaro-Winkler fuzzy matching and phonetic encoding to find duplicates before they reach the target. Composite scoring across email, name, phone, and address with configurable confidence thresholds.

"We don't know what state our Salesforce org is in."

Run 45 audit checks before the migration starts. DuckDB-powered analytics across security config, data quality, schema health, API usage, and cloud-specific rules. Get a clear report with findings and severity. Know the org before you touch it.

"The migration worked in sandbox but failed in production."

Five quality gates catch issues between extraction and loading. Validation rules, required fields, picklist mismatches, field length violations. If something will fail in production, it fails at the gate. Not at 2am.

"We spent two weeks mapping fields in a spreadsheet."

AutoMapper suggests mappings and accelerators give you a starting template to adjust. 97 transform functions handle format conversion, concatenation, date parsing, and validation. Lookup tables handle value mapping. Live preview shows results before execution.

What else is in here.

DuckDB audit engine, 45 checks, 14 cloud types, CLI binary
Adaptive batch sizing, adjusts based on API latency and error rate
Checkpoint/resume, sub-batch granularity with MD5 checksums
28 quality rules, email, phone, zip, amount, date validation
Migration health score, 0 to 100 go/no-go before every run
Schema drift detection, catches field changes mid-migration
Two-pass upsert, creates records first, resolves lookups second
Salesforce Bulk API v2: auto-selected for high-volume loads
Compliance reports, auto-generated SOC 2 and GDPR
Circuit breaker, stops cascading connector failures
Rate limiter, proactive throttling from API response headers
PII detection + masking, SSN and credit card patterns, redaction on export
Event log, 19 event types, JSON-lines, full record replay
Data lineage, per-field transformation tracking
Data catalog, search any field across all migrations
YAML connectors, define new sources without writing code
Incremental sync, only migrate records that changed
Identity resolution, Jaro-Winkler + Double Metaphone
10 conflict types, 9 resolution strategies
Expression language, Pratt parser, 38 functions, aggregations

Your data stays on your machine.

Pathfinder is a native macOS desktop application. Credentials live in the Keychain. Data lives in local SQLite and DuckDB. Processing is local. No cloud. No telemetry. No third-party access.

Keychain
credentials
100% local
processing
PII detection
& masking
Full audit
trail

Common questions.

What systems does it connect to?

Eight built-in connectors: Salesforce, Dynamics 365, HubSpot (target only), CSV, Excel, JSON, XML, and YAML-defined custom connectors for anything with an API.

How does the audit engine work?

Connect to a Salesforce org. Pathfinder extracts metadata, loads it into DuckDB, and runs 45 checks covering security configuration, data quality, schema health, API limits, and cloud-specific rules. Results export as JSON, Markdown, or HTML. Runs from the desktop app or the pf-audit CLI.

What happens when a record fails during migration?

For the eight most common Salesforce rejection types (REQUIRED_FIELD_MISSING, STRING_TOO_LONG, DUPLICATE_VALUE, rate limits, ENTITY_IS_DELETED, MALFORMED_ID, UNABLE_TO_LOCK_ROW), Pathfinder remediates and retries automatically. For everything else, it logs the error and continues. You configure the threshold.

How fast is it for large datasets?

Pathfinder auto-selects Salesforce Bulk API v2 for batches over 10,000 records. Adaptive batch sizing adjusts in real time. Checkpoint/resume means progress is never lost.

Is this a cloud service?

No. Native macOS desktop application. All processing is local. Credentials are in the macOS Keychain. The audit engine uses DuckDB locally. The only network traffic is between your machine and the systems you connect to.

Audit it. Migrate it. Trust it.

Pathfinder is in early access for consultants and implementation partners. Built in Rust. 8,931 tests, all passing. Ready for production data.

Discuss your migration