#4241·career-ops

[Source] Level (jobsbylevel.com): operator submission, public XML feed, measured under the five rules

Author: finalburnerCreated Sep 16, 2026Updated Sep 16, 2026

I operate this board, so I'm disclosing that first. Every figure below was measured against the live feed on 2026-09-16, and each one comes with a command you can rerun. The Jobgether/undisclosed-employer finding from the previous measurement is resolved now that the operator excludes those rows from the feed (see rule 1); the one finding still working against this proposal is feed size, which I've listed rather than leaving for review to find.

Source name and URL

Level: https://jobsbylevel.com. It's a job board for roles that involve AI work, and each posting carries a level marker from 1 to 4 for how much the role involves AI. The feed documentation is at https://jobsbylevel.com/feeds, and the level method is at https://jobsbylevel.com/method.

Who operates the source?

Level (jobsbylevel.com). I'm the founder and operator, so this is an operator submission. Level is also a commercial job board: employers can buy paid postings and boosts on the website. See rule 3 below for why none of that reaches the feed.

Operator contact at the source's own domain

[email protected], a role address at the source's own domain. If you'd like further proof of domain control, I can serve a file or add a DNS TXT record on jobsbylevel.com on request.

Public read access

These are static XML files. No auth, API key, cookie or session is needed.

File Content Measured 2026-09-16
https://jobsbylevel.com/feeds/jobs/1.xml/16.xml Full live inventory, split into 16 parts 21,556 jobs total (16-part sum, 09:13–09:14 UTC), about 11.6 MB per part, 186.1 MB in total (195,181,442 bytes; descriptions included). A second 16-part pass minutes later (~09:17 UTC, used for the rule 1/2 percentages below) totalled 21,552
https://jobsbylevel.com/feeds/jobs.xml The same full inventory as a single file 21,557 jobs, 195,170,773 bytes (186.1 MB), read separately at 09:18 UTC
https://jobsbylevel.com/feeds/jobs/recent.xml Jobs from the last 7 days only (part=recent-7d) 1,919 jobs
https://jobsbylevel.com/feeds/jobs/trovit/* The same data in Trovit format not used by the provider
https://jobsbylevel.com/feed.xml RSS not used by the provider

Each <job id="…"> element contains id, title, url (the Level page), link, company, company_url, apply_url (the employer's application URL), level (1–4), level_label, location, country, remote, category, contract_type, date, salary_min/salary_max/salary_currency/salary_frequency and description. Each file ends with an XML comment that reports its counts and exclusions, for example (part 1, measured 2026-09-16):

<!-- part=1/16 published=1299 no_country=57 short_description=0 no_title=0 undisclosed_employer=211 generated=2026-09-16T09:13:26.162Z -->

The footer now carries an undisclosed_employer= counter: rows whose employer isn't identifiable (the "listed on behalf of a partner company" rows from the last measurement) are excluded from published and counted there instead. That counter climbs between separate fetches taken minutes apart even when published barely moves (recent.xml: 3,486 → 3,556 → 3,779 across three fetches between 09:13 and 09:20; the 16-part sum read 3,823 at 09:13–09:14 while the single-file footer read 4,118 at 09:18), so it behaves like a running site-wide count rather than a fixed per-file total — treat the exact value as evidence the exclusion is happening, not as a precise per-scan denominator.

The parts are numbered, and the 16-part sum is close to the single-file total, though not identical between separate fetches: 21,556 at 09:13–09:14 UTC, 21,552 on a second 16-part pass around 09:17 UTC, and 21,557 from the single-file endpoint at 09:18 UTC. The board regenerates the feed on every request, so counts drift by a handful of jobs between fetches. A provider can still read the whole inventory with a bounded number of requests. /feeds/jobs/17.xml returned 404 on every pass.

Access by user agent, measured 2026-09-16: career-ops' DEFAULT_USER_AGENT (user-agent.mjs), curl/8.7.1 and node all get HTTP 200. Python's default Python-urllib/* user agent gets HTTP 403, which is why the reproduce scripts below set the career-ops user agent explicitly.

https://jobsbylevel.com/robots.txt now carries a dedicated block for the provider, measured 2026-09-16:

User-Agent: career-ops
Allow: /
Allow: /feeds/
Disallow: /api/
Disallow: /jobs/edit/
Disallow: /confirm/
Disallow: /unsubscribe/
Disallow: /panel/
Disallow: /go/
Disallow: /md/
Disallow: /*?q=
Disallow: /*&q=
Disallow: /*?category=
Disallow: /*&category=
Disallow: /*?tool=
Disallow: /*&tool=
Disallow: /*?remote=
Disallow: /*&remote=
Disallow: /*?seniority=
Disallow: /*&seniority=
Disallow: /*?minSalary=
Disallow: /*&minSalary=
Disallow: /*?level=
Disallow: /*&level=
Disallow: /*?sort=
Disallow: /*&sort=

Sitemap: https://jobsbylevel.com/sitemap.xml

That's one of 22 User-Agent blocks in the file. The other blocks, including *, still carry Disallow: /feeds/: the feed is opened to career-ops by name, not to every crawler. Since the provider sends the career-ops user agent, this dedicated block is the one that applies to it, and Allow: /feeds/ is explicit.

Rule 1 — real, attributed, free for candidates

  • Free: job pages are server-rendered and return HTTP 200 without an account, cookie or payment. Sample: https://jobsbylevel.com/jobs/product-marketing-lead-billing-at-stripe-488819. Candidates never pay, and the page's apply button reaches the employer through a single redirect (see rule 2).
  • Attributed: the full feed covers 352 distinct companies, confirmed again 2026-09-16 (the most frequent include OpenAI, Stripe, Databricks, Zscaler, Robinhood, Anthropic and Elastic).
  • Resolved since the last measurement: undisclosed-employer rows are now excluded from the feed. At the last measurement, 3,936 of 25,385 jobs (15.5%) had company = Jobgether, rising to 3,578 of 5,383 (66.5%) in recent.xml, and 3,823 of those descriptions said the position was "listed on behalf of a partner company" without identifying the actual employer. Measured again 2026-09-16, those undisclosed-employer rows are excluded from published and counted separately in each footer's undisclosed_employer= counter instead (see "Public read access"). The company = Jobgether rows that remain are down to 113 of 21,552 jobs (0.52%) in the full inventory and 91 of 1,919 (4.74%) in recent.xml, and none of the remaining ones contain the "on behalf of a partner company" wording (checked by the same regex as before, against 0 matches this time). I don't think Jobgether rows need a rule-1 carve-out any more; if you still find a remaining row that doesn't name a real employer, I'll remove it on Level's side.

Reproduce:

python3 - <<'EOF'
import urllib.request, xml.etree.ElementTree as ET, re
tot = jg = behalf = 0; companies = set()
UA = {"User-Agent": "Mozilla/5.0 (compatible; career-ops/1.0; +https://github.com/career-ops-hq/career-ops)"}
for i in range(1, 17):
    r = urllib.request.urlopen(urllib.request.Request(f"https://jobsbylevel.com/feeds/jobs/{i}.xml", headers=UA), timeout=120)
    for _, el in ET.iterparse(r):
        if el.tag != "job": continue
        tot += 1; c = (el.findtext("company") or "").strip(); companies.add(c)
        if c == "Jobgether":
            jg += 1
            behalf += bool(re.search(r"on behalf of a partner company", el.findtext("description") or "", re.I))
        el.clear()
print(tot, len(companies), jg, behalf)   # 2026-09-16 (~09:17 UTC pass): 21552 352 113 0
EOF

Rule 2 — canonical URL

  • apply_url is the employer's application URL, and url is the Level page, which can serve as secondary attribution. This is the same shape as remotli.mjs, and it's what providers/ADDING_A_PROVIDER.md (section 1) says Job.url should prefer.
  • Full inventory, measured 2026-09-16 (~09:17 UTC pass, 21,552 jobs): 21,295 of 21,552 jobs (98.81%) have an apply_url. Every one uses https:, and none points back to jobsbylevel.com. For the 257 jobs without one, the provider would fall back to the Level page.
  • The 7-day file, measured 2026-09-16 (09:20 UTC, 1,919 jobs): 1,890 of 1,919 (98.49%), with the same results (all https, zero self-links).
  • Most common apply hosts in the full inventory: jobs.ashbyhq.com 8,639, job-boards.greenhouse.io 5,380, jobs.lever.co 909, databricks.com 877, stripe.com 604, boards.greenhouse.io 599. In the 7-day file: jobs.ashbyhq.com 703, job-boards.greenhouse.io 467, jobs.lever.co 143, boards.greenhouse.io 80, stripe.com 70, jobs.smartrecruiters.com 51. jobs.lever.co fell from the top apply host (4,732 links, mostly jobs.lever.co/jobgether/…) to fifth, now that undisclosed-employer rows are excluded from the feed (see rule 1).
  • On the website, the apply button goes through /go/{id}, which is a single 302 to the same apply_url (checked at the last measurement: /go/000e89a1-…https://jobs.lever.co/jobgether/031440a5-…/apply; not rechecked in this pass). The provider never uses /go/, because the feed already carries the direct link.

Reproduce:

python3 - <<'EOF'
import urllib.request, xml.etree.ElementTree as ET, urllib.parse, collections
n = has = https = self_ = 0; hosts = collections.Counter()
UA = {"User-Agent": "Mozilla/5.0 (compatible; career-ops/1.0; +https://github.com/career-ops-hq/career-ops)"}
for i in range(1, 17):
    r = urllib.request.urlopen(urllib.request.Request(f"https://jobsbylevel.com/feeds/jobs/{i}.xml", headers=UA), timeout=120)
    for _, el in ET.iterparse(r):
        if el.tag != "job": continue
        n += 1; a = (el.findtext("apply_url") or "").strip()
        if a:
            has += 1; p = urllib.parse.urlparse(a); hosts[p.hostname] += 1
            https += p.scheme == "https"; self_ += (p.hostname or "").endswith("jobsbylevel.com")
        el.clear()
print(n, has, https, self_, hosts.most_common(6))
EOF

Rule 3 — promoted content and complete inventory

  • Paid placement exists on the site, but not in the feed. Level sells paid postings and boosts on jobsbylevel.com. The feed has no promoted, sponsored, featured or boost field. The measured tag set is exactly the one listed under "Public read access", and none of those names matches paid|boost|featur|promot|sponsor.
  • Ordering: every file is sorted by job id only. Checked 2026-09-16 on the 7-day file and across the 16 parts, read in order: the id sequence is sorted ascending as strings in both (the single-file endpoint was checked for job count and footer only in this pass, not re-checked for ordering). The ordering rule is published under "How jobs are ordered, and what cannot buy a place" at https://jobsbylevel.com/feeds. It says paid postings and boosts neither move a job in these files nor add one. Ranking stays on the user's machine, as rule 3 requires.
  • Complete inventory: the numbered parts are the full catalogue of live jobs, not a selection. The exclusions are expired jobs, jobs that fail a completeness check, and now undisclosed-employer/intermediary rows, and the footer comment of each file counts those exclusions. The full-inventory footer (single file, read 2026-09-16 09:18 UTC) reads published=21557 no_country=1079 short_description=3 no_title=0 undisclosed_employer=4118, and the 7-day file (09:20 UTC) reads published=1919 no_country=80 short_description=1 no_title=0 undisclosed_employer=3779. The provider would read all 16 parts, not recent.xml.
  • Finding against me: size. The full inventory is about 186 MB per scan (16 × ~11.6 MB), because descriptions are included — down from about 203 MB at the last measurement now that undisclosed-employer rows are excluded, but still the same shape of problem. If that's too heavy for a zero-token scanner, I can publish a lighter set of parts without description, still ordered by id and still complete, under a path you choose. Just let me know.

Reproduce (ordering, counts, footer):

python3 - <<'EOF'
import urllib.request, re
tot = 0; ids = []
UA = {"User-Agent": "Mozilla/5.0 (compatible; career-ops/1.0; +https://github.com/career-ops-hq/career-ops)"}
for i in range(1, 17):
    b = urllib.request.urlopen(urllib.request.Request(f"https://jobsbylevel.com/feeds/jobs/{i}.xml", headers=UA), timeout=120).read()
    tot += b.count(b"</job>"); ids += re.findall(rb'<job id="([^"]+)"', b)
    print(i, re.search(rb"<!--(.*?)-->\s*</jobs>", b).group(1).decode().strip())
print("total", tot, "sorted by id:", ids == sorted(ids))
EOF

Provider PR (optional)

None yet. Because this is an operator-run board, I'm following providers/ADDING_A_PROVIDER.md and opening this proposal before writing code. The Jobgether/undisclosed-employer question from the last measurement is resolved on the feed side; once you've decided how the payload size should be handled, I'll send a providers/level.mjs PR for job_boards:. It would use an explicit-only detect(), a fixed host, redirect: 'error', all 16 parts, apply_url first with the Level page as fallback, and fixture tests.

Anything else

  • Operator statement (rule 4): I run Level. Listing it isn't an endorsement, and I understand that no placement, traffic or permanence is owed. I'm fine with the operator being declared in SUPPORTED_JOB_BOARDS.md.
  • Rule 5 (single source): a provider would read only jobsbylevel.com's own feed and would do no cross-source work. Level reads each company's own public job board directly on Greenhouse, Lever, Ashby and SmartRecruiters (one board per company, apply_url points to that board) and doesn't scrape or republish other job sites. One of those boards belongs to an intermediary, Jobgether (a Lever board), which posts on behalf of unnamed partner companies; the feed now excludes those undisclosed-employer rows (see rule 1).
  • What the source adds: each job has a level marker (level 1–4 plus level_label) showing how much the role involves AI work. In the 7-day file, measured 2026-09-16, the counts are L1 1,429, L2 150, L3 251 and L4 89 (of 1,919 jobs). The method is at https://jobsbylevel.com/method. Operator-reported classifier accuracy, measured against a labelled sample of 6,487 examples labelled by Gemini: 74.1% exact match and 96.6% within one level. This is a data field, not a ranking signal. career-ops can ignore it or use it locally.
  • Volume and freshness: 21,552-21,557 live jobs (three separate fetches 2026-09-16, see "Public read access") across 352 companies. The feed regenerates on every request (generated= in each footer). Most jobs are remote or international.
  • Contact: [email protected], or in this thread.