← Back to PavedIT

Data provenance

How PavedIT builds its career graph

Beta notice · Updated August 8, 2026

Jobs

PavedIT prioritizes employer-authorized feeds, then official public ATS endpoints, then approved structured-data collection. Licensed providers are used only for measured coverage gaps. A source must pass terms review, robots policy where applicable, rate limits, canonical-URL validation, attribution, freshness, duplicate, and removal checks before its jobs can be published.

Career taxonomy

PavedIT owns its graph schema, mappings, corrections, and derived evidence—not the public classifications it builds upon. Production includes the versioned O*NET 30.3 occupation taxonomy, job-title aliases, essential-skill ratings, sample sizes, confidence bounds, dates, CC BY 4.0 license, and attribution. ESCO 1.2.1, SOC/BLS, ISCO, and future licensed releases retain their own source URI, version, license, and attribution when added.

Classification, crosswalks, occupation–skill requirements, observed career transitions, and task-level AI impacts are stored as different evidence types. A transition or AI-impact claim must include its method, release or model version, sample size, observation window, and uncertainty. Sparse evidence is withheld instead of replaced with a made-up score. Public evaluation records measure coverage, mapping precision and recall, hierarchy consistency, stability, calibration, language parity, and occupation parity.

Collection boundaries

PavedIT never bypasses logins, CAPTCHAs, paywalls, access tokens, IP blocks, or other technical restrictions. It does not collect candidate profiles. Approved crawlers identify themselves, provide a monitored contact address, stay within approved hosts, and can be disabled source by source.

Crawler identification

Approved automated collection identifies itself as PavedITBot/1.0 and carries a monitored contact address plus a link to this page in every request's User-Agent. It honors robots.txt where the source's policy record requires it, stays on each source's approved host within a recorded request budget, and follows only HTTPS targets. Site operators can request review or exclusion through the contact address in the crawler's User-Agent; each source can be deactivated individually and immediately.

Current live sources

Job and source totals change continuously, so PavedIT does not freeze a marketing count on this page. The live panel below reads the production health snapshot directly from Supabase, and each displayed listing must retain an attributable HTTPS employer-controlled application URL. PavedIT does not accept applications on the employer's behalf.

Loading the current catalog health snapshot…

The collector records source-level acceptance, duplicate, error, freshness, and removal evidence. New boards enter a private candidate registry, pass a fresh live-endpoint and HTTPS apply-link validation, and can be promoted only under an active, unexpired provider-rights review. An empty or smaller catalog remains preferable to unapproved or fabricated data.