Skip to main content

Our Methodology

How we research, collect, clean, and deliver custom datasets.

From source discovery to delivery

Every dataset follows the same nine-step process. No shortcuts. No black boxes.

Source Discovery

We research available public sources relevant to your request: open data portals, government registries, municipal databases, business directories, licensing bodies, published reports, and statistical agency releases. We evaluate whether the data exists, how accessible it is, how complete the records are, and whether collection is both lawful and practical. Sources are documented and included in delivery.

Feasibility Assessment

Before any work begins, we assess whether the requested dataset is actually buildable: are sufficient sources available? Are records reasonably complete? Is the scope realistic within your budget? Is collection lawful given the nature of the data and how it would be used? We communicate the results of this assessment honestly in our written proposal. If a dataset cannot be built as requested, we say so before accepting the project.

Data Acquisition

We collect data from identified sources using appropriate methods for each source type: direct download of published files, API access where available, structured extraction from public web pages, or manual extraction where records are published in formats that require it. We do not collect personal information, and we do not access private or access-controlled systems. Acquisition methods are documented and included in the source list delivered with your dataset.

Data Normalization

Raw collected data is cleaned and standardized. Column names are made consistent across sources. Data types are enforced: dates become dates, numbers become numbers, text fields are trimmed and consistent. Character encoding issues are corrected. Empty, null, and placeholder values are handled in a documented and consistent way. Structural inconsistencies between records from different sources are resolved.

Deduplication

Duplicate records are identified and resolved. When the same entity appears in multiple sources — or multiple times within a single source — records are compared and either merged into a canonical record or flagged for client review. Deduplication logic is documented: we note which fields were compared, what threshold was used, and how conflicts were resolved when source records disagreed.

Field Standardization

Field values are normalized to consistent formats across the dataset. Addresses are parsed into components and cleaned against standard formats. Phone numbers are formatted consistently. Postal codes are validated. Category and classification values from different sources are mapped to a unified taxonomy. Business names and organization names are standardized where variation is due to formatting rather than meaningful difference. All standardization decisions are documented in the data dictionary.

Geographic Enrichment

Where applicable and within scope, location data is enriched with additional geographic identifiers to support downstream analysis and mapping. This may include latitude and longitude coordinates, census dissemination areas, census subdivisions, provincial or territorial codes, electoral district boundaries, or other geographic identifiers relevant to your use case. Geographic enrichment is performed using authoritative public reference data and documented in the data dictionary.

Quality Review

The completed dataset is reviewed for coverage, accuracy, internal consistency, and alignment with the approved scope. We check for systematic gaps, outlier values, formatting regressions, and field completeness. A quality assurance report is produced documenting overall record counts, field-level completeness rates, any known limitations, caveats about source data quality, and any records that were excluded with a reason why. We do not deliver datasets with undisclosed quality issues.

Delivery & Documentation

We deliver the completed dataset in your requested format or formats, along with a full documentation package: a data dictionary defining each field name, data type, description, and example values; a source list naming every source used with collection dates and access methods; and where applicable, a quality assurance report. Delivery is via secure file transfer. We do not retain your dataset after delivery without your explicit consent.

Delivery formats

We deliver in the format you need. Multiple formats can be provided for a single dataset.

CSV

Comma-separated values — compatible with any spreadsheet application, database, or data analysis tool. UTF-8 encoded. The default delivery format for most datasets.

Excel

XLSX format with structured worksheets. Useful when the dataset includes multiple related tables or when the recipient's workflow is spreadsheet-based.

JSON

Structured JSON for use in applications, APIs, or data pipelines. Can be delivered as a flat array or as a nested structure matching your schema requirements.

GIS / Spatial

GeoJSON or Shapefile for datasets with geographic enrichment. Compatible with QGIS, ArcGIS, and other spatial analysis platforms. Available when geographic enrichment is included in scope.

Data Dictionary

Included with every delivery. Defines each field: name, data type, description, example values, and any applicable notes. Delivered as a separate document alongside the dataset.

QA Report & Source List

Every delivery includes a source list naming each data source used and a quality assurance report documenting record counts, field completeness, known limitations, and any records excluded with reasons.

A note on feasibility

Some datasets may not be available, complete, lawful to collect, or practical to build within a given scope or budget. Public data coverage varies significantly by jurisdiction, sector, and record type. Data that exists in one province may not exist in another. Records that are published at the federal level may not be granular enough for municipal use.

Each request is reviewed individually before a proposal is prepared. We will be upfront about what is and is not buildable. If your requested dataset is not feasible as described, we will explain why and suggest alternatives where possible.

Request a custom dataset

Describe what you need and we'll apply this methodology to assess feasibility and build your dataset. Submit a request to receive a written proposal with no obligation to proceed.