Philippines staffing research ·
How Should Duplicate Candidate Representation Be Researched?
A controlled study of duplicate submissions, match hypotheses, channel provenance, candidate contact, pause controls, and owner decisions.

Research question. When two sourcing channels present the same candidate, what evidence lets a hiring owner resolve the conflict without allowing a coordinator to decide candidate ownership, fees, or selection? This study treats duplicate representation as a provenance problem. It does not assume that the first email, first database timestamp, or most complete profile creates a valid commercial claim.
Why the question matters. A duplicate can arrive through an employee referral, an agency, a job board, a direct application, or a dormant talent record. Silently merging records can destroy consent and channel history. Leaving every record active can produce contradictory messages and count one person several times. The useful outcome is a traceable case that the authorized owner can decide, not an automated winner.
Unit and sample. Use one fictional person token and every observed presentation of that token to one requisition. Construct 120 cases across direct applications, referrals, two agencies, a talent community, a reopened requisition, and a confidential replacement. Include transliterated names, changed emails, shared phone numbers, an agency submission after a direct application, simultaneous events, and earlier interest in another role. All identities and documents remain synthetic.
Capture each presentation as its own event. Required fields are channel, sender, received instant, claimed event instant, requisition identifier and version, stated permission, attached terms, contact details, communications already sent, and the information visible at that time. Preserve the submitted artifact and hash. Do not copy a later interpretation into an earlier event or collapse unknown, not supplied, disputed, and not applicable into blank.
Identity matching remains a hypothesis. Compare stable tokens where available, then apply an approved rule to combinations of contact points and candidate-confirmed context. A similar name must not trigger a merge. A matching address can still represent a shared or recycled mailbox. Record probable duplicates separately until an authorized reviewer accepts or rejects the relationship, and retain the evidence that supported the original classification.
Build a provenance graph instead of a winner column. Nodes represent submissions, consent events, requisition versions, pauses, and communications. Edges state why records may describe one person or relate to one search. The graph must expose missing consent, an unknown requisition, expired commercial terms, or an assertion that cannot be corroborated. Restrict the view to fields needed for the declared review purpose.
Sequence test. Replay the queue by received time and again by verifiable event time. An agency may upload first while a direct application happened earlier; a referral may be logged late; a scheduled integration may create an apparently old record. Measure how often “earliest system timestamp” disagrees with the complete chronology, whether outreach stops while facts are disputed, and whether the responsible owner receives the unresolved alternatives.
Purpose test. For every node, record the declared purpose and candidate-facing evidence supporting it. Do not infer permission to submit someone to a new opening from a prior conversation about another job. The Philippine Data Privacy Act’s transparency, legitimate-purpose, proportionality, accuracy, currency, security, and retention principles frame the controls. They do not settle a representation dispute or prove a particular submission lawful.
Decision boundary. Coordinators may detect likely duplication, pause parallel outreach under an approved rule, assemble chronology, ask authorized factual questions, and route the case. Hiring and commercial owners decide representation, contractual priority, fees, channel exceptions, and next contact. Count a coordinator-made ownership ruling as a serious boundary failure even when the owner later reaches the same result. Correct guesses are still unauthorized decisions.
Candidate-experience trial. Seed one case where three channels contact the person within an hour, one where a pause reaches only an agency, and one where a merge hides the preferred contact route. Observe message count, contradictory instructions, response burden, correction time, and time to a single accountable contact. Faster internal closure is not success if duplicate or misleading outreach continues.
Controlled comparison. Process half the cases with a timestamp-ordered deduplication queue and half with the provenance graph plus pause state. Keep staff, information, and time equal. Compare detected duplicates, false merges, parallel outreach after detection, unsupported ownership rulings, missing acknowledgments, owner response time, and candidate-facing corrections. Report counts and denominators; do not let a favorable average hide a wrongful merge.
Adversarial cases include two different people with the same name, one person using a preferred name, a referral copied from a public profile without confirmed contact, an agency claim lacking a requisition, and a direct application received while the opening was paused. Add a merged record later proven false. The procedure must unwind that merge without rewriting the original trail or exposing one person’s record to another.
The owner packet contains a compact chronology, identity-match confidence and basis, purpose and permission evidence, applicable channel terms supplied by their owners, communications already sent, unresolved facts, and the next reversible action. It must not rank the candidate or recommend which commercial claim to honor. A separate event captures the accountable owner’s decision, rationale, scope, and recipients.
Privacy review inventories views, downloads, notifications, search indexes, integration caches, and deletion or correction paths. Candidate identifiers stay out of general dashboards. Temporary access expires. Rejected match hypotheses receive a defined retention treatment. If unexpected real data enters the exercise, testing stops and follows the approved incident route rather than copying the material into a defect report.
Interpretation separates observations, match classifications, owner decisions, and researcher inference. A high detection rate is useful only beside false-merge and excess-contact rates. A low conflict count could mean good routing, missed detection, or suppressed reporting. When evidence cannot establish identity, sequence, or permission, the defensible result is uncertainty with a named owner and a stop condition.
Limitations. Synthetic cases cannot establish a candidate’s intent, interpret agency contracts, authorize processing, or decide selection. They can show whether a proposed coordination service preserves facts and respects the decision boundary. A buyer should request a sanitized duplicate register, pause control, candidate correction path, access log, and owner acknowledgment, then begin with a narrow approved pilot.
Sources checked October 2, 2026: National Privacy Commission, “Republic Act 10173 — Data Privacy Act of 2012,” https://privacy.gov.ph/data-privacy-act/; National Privacy Commission, “The Data Privacy Act and Its IRR,” https://privacy.gov.ph/the-data-privacy-act-and-its-irr/; NIST, “Digital Identity Guidelines: Identity Proofing and Enrollment,” https://pages.nist.gov/800-63-4/sp800-63a/proofing/. NIST supplies technical concepts, not Philippine employment rules, and none of these sources proves provider performance.
Calibration for duplicate-candidate provenance. Two reviewers independently handle unseen edge cases and cite the exact evidence behind each classification. Disagreement becomes a finding about definitions, access, or source quality rather than a training score. Preserve both attempts, let the accountable owner clarify the rule, version that change, and retest with new cases so familiarity cannot masquerade as repeatability.
Recovery challenge for duplicate-candidate provenance. Remove a required source, delay an acknowledgment, replay an obsolete event, and interrupt the primary system. Observe whether the routine preserves last-known state, prevents double action, exposes uncertainty, and resumes without rewriting history. Record affected downstream copies and named recovery owners until each is verified or honestly remains unresolved.
Evidence review for duplicate-candidate provenance. Trace every sampled outcome backward to its source and forward to recipients. Distinguish observed fact, rule classification, researcher inference, and owner decision. Missing evidence stays missing. Measure coverage, exception age, access scope, propagation, and agreement with explicit denominators, while reporting serious boundary failures outside any aggregate score.
Acceptance for duplicate-candidate provenance. Set thresholds before opening the hidden answer key, including zero tolerance for unauthorized substantive decisions and avoidable sensitive-data exposure. A corrected outcome does not erase its first-pass failure. Change the control through its owner, retain the initial record, and demonstrate improvement only on a fresh blinded sample that includes adverse cases.
Procurement use for duplicate-candidate provenance. Ask the provider to demonstrate a sanitized register, version history, permission view, exception route, recipient acknowledgment, and audit export. A polished demo or policy is point-in-time evidence, not proof of continuing operation. Begin live work with a narrow approved population, least privilege, named reviewers, monitored exceptions, and a stop rule.
Decision brief for duplicate-candidate provenance. Present the buyer with the observed result, denominator, excluded cases, uncertainty, operational consequence, control cost, and accountable next decision. Include the strongest alternative explanation and the evidence that supports or weakens it. Avoid a single maturity label that blends boundary violations with ordinary delays. A useful recommendation identifies what can be delegated now, what must remain with the owner, which evidence is absent, and the next bounded test. The conclusion expires when a material source, system path, role boundary, or process version changes, so record the applicable scope and review trigger.
Limit the claim for duplicate-candidate provenance to the tested population, systems, versions, recipients, and observation window. Describe exclusions and cases that could not be determined. A passing result supports a cautious pilot decision; it does not guarantee future performance, legal compliance, security, employee outcomes, or accuracy outside the sample. Schedule review when ownership, source definitions, integrations, or access patterns change.