How a posting gets here, and how it leaves
1. Where postings come from
Roles are read from the employer's own publication, in this order of preference:
- Applicant tracking systems. The board a company actually hires through - Greenhouse, Lever, Workday, Ashby, SmartRecruiters, Recruitee, Personio, Teamtailor, JazzHR, Huntflow, Talentera and dozens more. The board is tied to the employer only after the ownership is confirmed from the company's own domain, so one vendor's page is never credited to the wrong company.
- Career pages. Where a company publishes without a vendor board, the page itself is read and parsed.
- Public feeds. A small number of aggregating feeds, used only where they carry the employer's own posting and a link back to it.
Every posting keeps the URL it was read from. A posting that will not say who is hiring is not listed at all: an employer with no name cannot be verified, and an anonymous listing is the single strongest marker of a re-published or speculative advert.
2. Re-checking, and what "live" means
The source board is refetched on a cadence tuned per employer, and a role counts as confirmed when the board it came from was successfully read within the last 48 hours and still listed it. The share of ATS-sourced roles meeting that test right now is 97%; it is recomputed from the database, not asserted.
When a board stops listing a role, the role is closed here and the moment is written into a lifecycle journal - opened, closed, reopened, expired, restored, reposted. That journal is what later scoring reads; nothing is inferred from a listing's age alone.
Postings that do not come from a vendor board expire after 60 days without a confirmation. ATS rows are deliberately exempt from that clock, because their timestamp is the employer's own publication date rather than a last-seen stamp - dating them by a rule that never applied to them is how a healthy listing ends up declaring an expiry in the past.
3. Stale and ghost listings
A role that stays open unusually long is not automatically fake - some employers simply hire slowly - so the threshold adapts to the employer rather than to the calendar:
- Stale when the role has been open longer than three times that employer's own median time to fill, and never below a floor of 90 days. The median is taken over a year of that company's closed roles and needs at least five of them; without a sample, only the floor applies.
- Repost when the same role has been taken down and put back up two or more times, counting only cycles at least a day apart.
- Company-wide staleness when at least half of an employer's open roles are stale and it has at least ten of them.
- Ghost only when staleness is joined by a second, independent signal - reposting or company-wide staleness - or when a role has been recycled three times or more. Staleness on its own is shown as a warning, never as a verdict.
A ghost listing is dropped from the sitemap, marked noindex and kept out of feeds and matches, while staying readable for anyone who follows a link to it. We would rather warn than delete: the listing may be genuine and merely slow.
4. Liveness: is anyone still hiring behind this?
A posting being online says nothing about whether a search is still running behind it, and that is the one thing a candidate cannot see. So every listed role carries a liveness score, shown beside its date as Hiring now, Likely open, Fading or Long shot with the number behind it. It is the product of three probabilities, each computed from the journal rather than asserted, and the popup on the chip names the three factors for that role:
- Open - the board still carries it. For a role read from an employer's applicant tracking system, the moment that board last confirmed the posting: within 48 hours counts in full, within 7 days at 0.9, within 14 days at 0.7, older at 0.4. A row the scanner wrote after the board's last full read is confirmed from the moment it was written, so a role posted this morning is never "verified yesterday". A posting past its own expiry date is closed, whatever the board says. A role from a career page or feed decays instead with the time since the source last showed it: 14 days in full, 30 at 0.85, 45 at 0.6, older at 0.4.
- Active - a search is running. From a year of closes in the lifecycle journal (closes younger than two days excluded as board flicker), a curve per cohort - seniority, company size, office or remote - gives the share of postings alive at each age that still close within the next 60 days. Only about 40% do even on day zero, because a third of every board never closes at all, so the curve is read as a decay against its own day-zero value rather than as a probability: a fresh posting starts at 0.86, and a 180-day posting keeps about a sixth of that. A curve speaks only with at least 200 alive postings in the cell; thinner cells roll up to seniority, then format, then the whole corpus. An employer with five or more closes of its own maps the age through its own median time to fill. The value is then multiplied by what the journal knows about this posting: reposted twice x0.75, three times or more x0.5, posted by an agency x0.6, a board where at least half the roles are stale x0.7, an evergreen pipeline posting (the same role recycled three times or more) x0.6, a placeholder title x0.5, part of a fresh hiring wave x1.1, an employer signalling urgency x1.1, a burst of new roles x1.05. A ghost verdict caps it at 0.3.
- Room - there is still space in the funnel. Where the posting stands in the employer's own median time to fill (its own with five or more closes, otherwise the cohort's): 1.0 up to 35% of the window, 0.9 to 70%, 0.75 to 100%, 0.55 to 150%, 0.45 to 200%, 0.35 beyond. On top of that an applicant-pressure index - days of life times the crowd a role draws, with worldwide-remote x2, junior x1.5, an employer with a hundred or more open roles x1.5, a top-decile reader count x1.5, a niche role x0.6 - takes a further x0.8 once it says the shortlist has had time to fill.
Score = 100 x Open x Active x Room. 70 and above is hiring now, 45-69 likely open, 25-44 fading, below 25 a long shot; an evergreen pipeline posting never reads above 40. The slow inputs - the close curves, duplicates, board staleness, reader attention - are rebuilt nightly, and the score is a pure function of them, so a posting that arrived after the pass is scored at render time from what its own row already knows. Right now 32% of listed roles read hiring now or likely open; the figure is recomputed from the database, not asserted.
Liveness is an employer-side estimate and is the same for every reader. It never enters a candidate's match score, never hides a role, and is calibrated against the journal: the share of postings in each band that closed within 60 days without reopening. Where employers report outcomes through the application loop, those replace the journal as the calibration source.
5. Employer truth index
The same journal scores an employer's board as a whole. An employer with at least three scored open roles carries a truth index, shown as Trust A to Trust E on its cards with the numbers behind it: 100 minus 60 times its ghost share, 25 times its stale share and 15 times its repost share, plus 5 where its median time to fill is under 21 days, minus 5 where it is 60 days or more. A from 85, B from 70, C from 55, D from 40, E below. It says how much of what this employer publishes turns into a real, finite search; it says nothing about the employer as a workplace.
6. Pay
Where an employer states a salary, that figure is shown as stated, converted to USD per year for comparison, with the original currency and period preserved. Ranges posted net of tax are grossed up using the local income-tax rate before comparison.
Where no salary is stated, an estimate is computed from the salaried part of the corpus and labelled an estimate everywhere it appears:
- Observations are postings with stated pay from the last 18 months, weighted by recency with a 180-day half-life, so a year-old posting votes with about a quarter of a fresh one.
- One employer contributes at most 8 effective rows to any cell, so a company that posts the same role fifty times cannot set the market rate.
- A cell - role, seniority and geography - speaks only with at least 8 effective observations; thinner cells lean on their parent cell, and a role outside every sufficiently sampled cell gets no estimate at all.
- The cell figure is then adjusted by residual factors for the employer, the city and the stack, each shrunk toward neutral and clamped, so a single outlier cannot move it far.
An estimate is display and sort information. It is never published as the employer's own baseSalary in structured data, never used to hard-fail a match, and never posted to a channel as though the employer had said it. In schema.org terms a stated figure is baseSalary; an estimate is estimatedSalary.
7. Roles, technologies and geography
Titles are noisy, so a posting is classified rather than trusted. Its title and body are read into a two-level taxonomy - a sphere such as Backend, Data Science or Hardware, then a specific role - with per-sphere gates that keep a posting out of a sphere it only mentions in passing. Technologies are matched against an indexed dictionary of 3,549 tools, frameworks and languages, and the stack that decides a match is the one the requirements ask for, not every word in the advert.
Location is resolved to real places: an office city where the employer names one, a country where it does not, and a remote scope - worldwide, country-restricted or region-restricted - where the work is remote. A remote role is not listed as an office role in a city it merely mentions.
8. What is deliberately kept out
- Postings that do not name the employer.
- Staffing and outsourcing agencies' re-publications of other companies' roles, unless a reader asks to see agency listings.
- Duplicates of a role already listed from the employer's own board.
- Personal contact details scraped from postings, which are held for the platform's own use and never published.
9. Corrections
If a role is listed that should not be, or a fact about your company is wrong, write to us - corrections are made at the source record so a rescan does not undo them. Personal-data requests go through the data subject request form.
The datasets this methodology produces, with sample files and a data dictionary, are described on the data page.
