# Scaffold People Data Labs enrichment app Defines the data model for enriching **Person** and **Company** with People Data Labs data. **Scaffold only** — the enrichment logic (the "mapper") follows separately; see the package README. ## Included - **Fields** on Person & Company (PDL base data set). - **Enums as SELECT / MULTI_SELECT** validated against PDL canonical files (v34.1). - **Standard-field mapping**: no `pdl*` shadow where a standard field exists. - **Location → ADDRESS**; **relation** `pdlCurrentCompany` ↔ `pdlCurrentEmployees`. - **Metadata**: `pdlId`, `pdlLikelihood`, `pdlEnrichmentStatus`, `pdlLastEnrichedAt`, `pdlRawPayload`. - Shared option constants + helper, indexes, and a view per object.
7.3 KiB
People Data Labs enrichment app
Enriches Person and Company records with People Data Labs (PDL) data.
Status: data-model scaffold. This package defines the fields, relation, indexes, views, role, and manifest. The enrichment logic function (the "mapper") is not yet implemented — see What the mapper must do.
Data-model decisions
Bundle scope
Only the core PDL company fields are defined. Premium / Comprehensive / specialized fields
(inferred_revenue, linkedin_follower_count, employee growth/churn/tenure, parent /
subsidiary, exec movement, top employers, funding_details, …) are out of scope for this app.
Enums → SELECT / MULTI_SELECT
Every PDL enum that has a canonical file is a SELECT, validated 0-missing/0-extra against PDL schema v34.1:
| Field | Type | Options |
|---|---|---|
pdlSeniority (job_title_levels, array) |
MULTI_SELECT | 10 |
pdlFundingStages (funding_stages, array) |
MULTI_SELECT | 29 |
pdlIndustry / pdlJobCompanyIndustry (industry) |
SELECT | 147 |
pdlJobTitleSubRole (job_title_sub_role) |
SELECT | 106 |
pdlJobTitleClass, pdlInferredSalary, pdlSex, pdlCompanyType, pdlSizeRange, pdlJobCompanySize, pdlLatestFundingStage, pdlLocationContinent, pdlLocationMetro, pdlMicExchange |
SELECT | 5 / 11 / 2 / 6 / 8 / 8 / 29 / 7 / 384 / 70 |
- Option
values are normalized to GraphQL enum names (united states→UNITED_STATES): uppercase, accents stripped, non-alphanumeric →_, digit-leading prefixed. - Option
universalIdentifiers are unique per field (shared enums like industry get a separate id-set per field). - Stays
TEXT(no canonical PDL enum file exists):pdlIndustryDetail(industry_v2),pdlJobOnetCode,pdlLocationRegion.
Standard-field mapping
pdl* shadows are removed where an equivalent standard field exists; the mapper writes
the standard field instead:
| Object | Removed shadow → standard target |
|---|---|
| Person | pdlLinkedinUrl→linkedinLink, pdlJobTitle→jobTitle, pdlFullName→name, pdlWorkEmail/pdlPersonalEmails→emails, pdlMobilePhone/pdlPhoneNumbers→phones |
| Company | pdlLinkedinUrl→linkedinLink, pdlWebsite→domainName, pdlDisplayName→name |
Shadows are kept where no reliable standard field is available: pdlEmployeeCount,
pdlTwitterUrl.
Trade-off: PDL's work/personal-email and mobile/other-phone distinction is dropped (folded
into the standard bags).
Location → ADDRESS composite
- Company location → the standard
addresscomposite (street/city/state/postcode/country/geo). - Person has no standard address field → dedicated
pdlLocation(ADDRESS). pdlLocationMetro(both) andpdlLocationContinent(company) stay SELECT — ADDRESS has no slot. Trade-off: ADDRESScountryis free text, so the country SELECT was dropped.
Relation
Dedicated pdlCurrentCompany (Person MANY_TO_ONE → Company) ↔ inverse
pdlCurrentEmployees (Company ONE_TO_MANY → Person). Deliberately not the standard
company relation, so PDL's detected employer can't overwrite the user's CRM account link.
Enrichment metadata
pdlId— PDL record id (re-enrich by id: more precise than by email).pdlLikelihood(Person, NUMBER) — PDL match confidence 1–10.pdlEnrichmentStatus(SELECT:MATCHED/NOT_FOUND/ERROR) — distinguishes "no match" from "never tried" (drives re-enrichment scheduling).pdlLastEnrichedAt(DATE_TIME),pdlRawPayload(RAW_JSON, full response).
Other
pdlTotalFundingisCURRENCY(mapper must convert the bare USD float → micros).- Indexes:
pdlIdandpdlLastEnrichedAton both objects. - Views: a curated "People Data Labs" TABLE view per object.
- Role: read/update on Person & Company (object-level; tighten to field-scoped later).
What the mapper must do
The logic function (to be built) must:
Orchestration
- Trigger via manual command-menu action / record create / batch (TBD).
- Call PDL Person and/or Company Enrichment with
PDL_API_KEY; passmin_likelihood/required_fieldsto control match quality. - On
200→ write fields + setpdlEnrichmentStatus = MATCHED; on404→NOT_FOUND; on error →ERROR. - Respect PDL rate limits (queue / throttle on
429). - TTL guard: skip re-enrichment if
pdlLastEnrichedAtis recent; prefer re-enriching bypdlId.
Field writing
- Write standard fields (fill-only-if-empty to avoid overwriting user data):
Person
name,emails,phones,linkedinLink,jobTitle; Companyname,domainName,linkedinLink,address. - Write
pdl*fields for everything else. - SELECT guard: only write a SELECT/MULTI_SELECT value if the normalized value exists in
the field's option set; otherwise skip and keep it in
pdlRawPayload(handles PDL schema versions newer than v34.1). Use the same normalization as the optionvalues. - MULTI_SELECT arrays:
job_title_levels→pdlSeniority;funding_stages→pdlFundingStages. - CURRENCY:
total_funding_raised(USD float) →{ amountMicros: value × 1_000_000, currencyCode: 'USD' }. - ADDRESS: split PDL
location.*into the composite — Company → standardaddress, Person →pdlLocation. - Relation: resolve
job_company_id→ find/upsert a Company record → linkpdlCurrentCompany. - Dates: handle partial PDL dates (
YYYY,YYYY-MM) forjob_start_date,last_funding_date,birth_date. - Always set
pdlId,pdlLastEnrichedAt,pdlRawPayload,pdlLikelihood(person).