← Documentation

Adapter and catalog contract

Dataset

Version 1 catalog JSON has schema_version, generated_at, opportunities and sources. Generation time describes the dataset build, not the freshness of every record. Consumers reject unsupported schema versions and malformed records before publication.

Opportunity fields

FieldMeaning
idStable source record identity
titleOriginal title, escaped when rendered
urlOriginal opportunity article or publisher detail page
sourceRegistry slug
source_urlRetrieval endpoint such as the RSS feed
summaryPlain-text excerpt, at most 600 characters
languageSource language; do not silently translate
categoryscholarships, internships, volunteering, fellowships, training, jobs, competitions, grants or other
tagsOriginal or explicitly mapped labels
host_countriesISO country codes where opportunity occurs; empty means unknown
eligible_countriesExplicit eligible applicant countries; empty means unknown
deadlineISO date/time or null; null never means open indefinitely
published_atPublisher publication time or null
first_seen_at / created_atFirst successful observation
last_seen_atMost recent successful observation of this record
last_checked_atSuccessful source retrieval time, not a failed attempt time
updated_atActual content update time
statusopen, expired or unknown; avoid unsupported certainty
locationOriginal location text or null

Publisher location must never be used as applicant eligibility. Preserve provenance when grouping duplicate opportunities; conservative URL identity is preferable to merging unrelated titles. HTTP and HTTPS links only; reject credentials, script protocols and invalid URLs.

Source registry

Each source has a safe slug, name, homepage, retrieval endpoint, adapter type, language, publisher country, enabled status and documentation. Research records include added/checked dates, types of content, evidence of recent activity, retrieval method and limitations. Directory candidates are distinct from tested live adapters. Do not mark sources active simply because the homepage exists.

Adapter responsibilities

See the Community adapter implementation guide for manifest fields, trusted dispatch, commands and reviewed programme selection.

A common RSS/Atom adapter handles configuration-only sources. A custom adapter transforms fetched source content into normalized candidates. The pipeline owns stable IDs, timestamps, common validation, persistence and publication. Fixture tests exercise typical and malformed content without hitting live services. Implementation function names must match the actual exported API; this contract describes responsibilities rather than requiring a specific signature.

Collection uses bounded response sizes, timeouts, restricted public endpoints and polite request frequency. Retry transient errors with a bounded budget. Source errors are isolated. Never overwrite good state with empty data caused by parse/network failures; successful empty responses require an explicit documented policy. Missing an item from a feed is not proof that the opportunity expired.

Freshness acceptance

A failed fetch must not parse old latest.xml as current data or reset first-seen dates. Retained records keep observation timestamps and gain source failure/stale context. Tests must show that failed collection cannot falsely refresh cards. Existing historical records with unreliable timestamps should carry that limitation rather than fabricated verification dates.