Our Methodology
How we research, collect, clean, and deliver custom datasets.
The Process
From source discovery to delivery
Every dataset follows the same nine-step process. No shortcuts. No black boxes.
Source Discovery
We research available public sources relevant to your request: open data portals, government registries, municipal databases, business directories, licensing bodies, published reports, and statistical agency releases. We evaluate whether the data exists, how accessible it is, how complete the records are, and whether collection is both lawful and practical. Sources are documented and included in delivery.
Feasibility Assessment
Before any work begins, we assess whether the requested dataset is actually buildable: are sufficient sources available? Are records reasonably complete? Is the scope realistic within your budget? Is collection lawful given the nature of the data and how it would be used? We communicate the results of this assessment honestly in our written proposal. If a dataset cannot be built as requested, we say so before accepting the project.
Data Acquisition
We collect data from identified sources using appropriate methods for each source type: direct download of published files, API access where available, structured extraction from public web pages, or manual extraction where records are published in formats that require it. We do not collect personal information, and we do not access private or access-controlled systems. Acquisition methods are documented and included in the source list delivered with your dataset.
Data Normalization
Raw collected data is cleaned and standardized. Column names are made consistent across sources. Data types are enforced: dates become dates, numbers become numbers, text fields are trimmed and consistent. Character encoding issues are corrected. Empty, null, and placeholder values are handled in a documented and consistent way. Structural inconsistencies between records from different sources are resolved.
Deduplication
Duplicate records are identified and resolved. When the same entity appears in multiple sources — or multiple times within a single source — records are compared and either merged into a canonical record or flagged for client review. Deduplication logic is documented: we note which fields were compared, what threshold was used, and how conflicts were resolved when source records disagreed.
Field Standardization
Field values are normalized to consistent formats across the dataset. Addresses are parsed into components and cleaned against standard formats. Phone numbers are formatted consistently. Postal codes are validated. Category and classification values from different sources are mapped to a unified taxonomy. Business names and organization names are standardized where variation is due to formatting rather than meaningful difference. All standardization decisions are documented in the data dictionary.
Geographic Enrichment
Where applicable and within scope, location data is enriched with additional geographic identifiers to support downstream analysis and mapping. This may include latitude and longitude coordinates, census dissemination areas, census subdivisions, provincial or territorial codes, electoral district boundaries, or other geographic identifiers relevant to your use case. Geographic enrichment is performed using authoritative public reference data and documented in the data dictionary.
Quality Review
The completed dataset is reviewed for coverage, accuracy, internal consistency, and alignment with the approved scope. We check for systematic gaps, outlier values, formatting regressions, and field completeness. A quality assurance report is produced documenting overall record counts, field-level completeness rates, any known limitations, caveats about source data quality, and any records that were excluded with a reason why. We do not deliver datasets with undisclosed quality issues.
Delivery & Documentation
We deliver the completed dataset in your requested format or formats, along with a full documentation package: a data dictionary defining each field name, data type, description, and example values; a source list naming every source used with collection dates and access methods; and where applicable, a quality assurance report. Delivery is via secure file transfer. We do not retain your dataset after delivery without your explicit consent.
Formats
Delivery formats
We deliver in the format you need. Multiple formats can be provided for a single dataset.
CSV
Comma-separated values — compatible with any spreadsheet application, database, or data analysis tool. UTF-8 encoded. The default delivery format for most datasets.
Excel
XLSX format with structured worksheets. Useful when the dataset includes multiple related tables or when the recipient's workflow is spreadsheet-based.
JSON
Structured JSON for use in applications, APIs, or data pipelines. Can be delivered as a flat array or as a nested structure matching your schema requirements.
GIS / Spatial
GeoJSON or Shapefile for datasets with geographic enrichment. Compatible with QGIS, ArcGIS, and other spatial analysis platforms. Available when geographic enrichment is included in scope.
Data Dictionary
Included with every delivery. Defines each field: name, data type, description, example values, and any applicable notes. Delivered as a separate document alongside the dataset.
QA Report & Source List
Every delivery includes a source list naming each data source used and a quality assurance report documenting record counts, field completeness, known limitations, and any records excluded with reasons.
A note on feasibility
Some datasets may not be available, complete, lawful to collect, or practical to build within a given scope or budget. Public data coverage varies significantly by jurisdiction, sector, and record type. Data that exists in one province may not exist in another. Records that are published at the federal level may not be granular enough for municipal use.
Each request is reviewed individually before a proposal is prepared. We will be upfront about what is and is not buildable. If your requested dataset is not feasible as described, we will explain why and suggest alternatives where possible.
Get started
Request a custom dataset
Describe what you need and we'll apply this methodology to assess feasibility and build your dataset. Submit a request to receive a written proposal with no obligation to proceed.