Changelog

Release history for the ByteKit SDKs, CLI, and MCP server

For AI agents: https://bytekit.com/llms.txt — every docs page is available as markdown by appending .md

Each section below is the CHANGELOG.md kept alongside that client’s source in the ByteKit repository. Entries record what a release changed, including names and options it retired.

@hunt-labs/bytekit-sdk

TypeScript SDK, published on npm.

[Unreleased]

[0.11.9] - 2026-10-02

Changed

  • Terms-version doc comments describe the new onboarding contract (#5055). Bootstrap no longer records Terms acceptance: the internal POST /v1/account/bootstrap terms_version body is now documented as a version guard, with acceptance recorded by POST /v1/account/accept-terms. The account terms_version description now says an empty string means the account has not accepted the Terms yet. Regenerated src/generated/api.d.ts; description text only, no type or SDK method signature moves.

[0.11.8] - 2026-09-30

Changed

  • Bulk job types add total_billed_bytes and deprecate the bulk credit fields (#5021). Every bulk job body (POST/GET/DELETE /v1/bulk, /v1/scrape/bulk, /v1/fetch/bulk, and the GET /v1/bulk list) gains total_billed_bytes: number | null: the bandwidth billed for the job, in bytes, after multipliers. It is null on the create response and for a bulk that predates the field. total_credits_charged and the bulk item credits_charged are marked @deprecated; they report 0 under the dual and bytes billing modes. Regenerated src/generated/api.d.ts; no SDK method signature moves.

[0.11.7] - 2026-09-29

Changed

  • Internal GET /v1/billing/topup-balance types add bandwidth.debt_bytes (#4960, #4975). The bytes the account still owes for bandwidth a finished request used beyond everything it could pay for, net of any repaid. Declared on that endpoint only, not on the shared BandwidthBalance schema GET /v1/usage returns. No SDK method calls it.

[0.11.6] - 2026-09-24

Changed

  • Regenerated src/generated/api.d.ts (#4865): a documentation-wording correction to the POST /v1/schema billing description, which now says the page bandwidth is billed at its upstream wire bytes with no endpoint factor instead of "exactly like /v1/scrape". No behavior change: doc comments only, no SDK signature moves.

[0.11.5] - 2026-09-22

Changed

  • Internal GET /v1/logs types list schema rows (#2373). Both endpoint enums add schema, for /v1/schema structured-extraction requests, whose row ids carry the new sch_ prefix. No SDK method calls it.

[0.11.4] - 2026-09-21

Changed

  • Regenerated src/generated/api.d.ts (#4732): the 422 responses of GET /v1/fetch and POST /v1/fetch declare the optional X-Fetch-ID header, present when the failed request was recorded and naming its row in Logs (sc_<hex> for the header's ft_<hex>). Doc comments only; no SDK signature moves.

[0.11.3] - 2026-09-21

Changed

  • Regenerated src/generated/api.d.ts (#4669): GET /v1/fetch and POST /v1/fetch document a second 422 condition. A fetch whose retries were exhausted against the target used to answer 503 with code internal_error and the message "Fetch service unavailable.", which read as an outage of the API; it now answers 422 with code blocked when the target's bot protection refused every attempt and upstream_error (message "Failed to fetch URL.") otherwise, and it is still charged nothing. A client that matched that 503 must match 422 and the code instead; the 503 internal_error remains only for a deployment whose fetch service endpoint or internal token is not configured. Doc comments only; no SDK signature moves.

[0.11.2] - 2026-09-18

Changed

  • Regenerated src/generated/api.d.ts (#4582): bulk item billing is BillingNode | null.
  • Regenerated src/generated/api.d.ts (#4584): POST /v1/schema now rejects slow-path rendering options (wait_until=networkidle, delay_ms, wait_for_selector, cookies, custom headers) with error code unsupported_option instead of unsupported_url, still HTTP 400. unsupported_url is removed from the API and has no other emitter; a client matching on it must match unsupported_option for this case. Doc comments only; no SDK signature moves.

[0.11.1] - 2026-09-17

Changed

  • Regenerated src/generated/api.d.ts (#4563): the internal listBulk operation's status query parameter and list-item status now admit cancelled. Additive; no SDK method calls listBulk, so no published signature moves.

[0.11.0] - 2026-09-14

Added

  • solve_challenge on scrape.create (#4447). ScrapeOpts now carries solve_challenge?: boolean, and the generated ScrapeRequest schema type carries the field with its server default false. When true, the scrape attempts to clear a Cloudflare challenge; success is not guaranteed. It is billed as a normal browser render. A scrape that omits the option sends no solve_challenge key, so existing calls are unchanged on the wire.
  • credits_scrape on the usage response types (#4452). The generated GET /v1/usage response type and each GET /v1/usage/daily data item now carry credits_scrape: number: credits drawn under the scrape usage source, such as /v1/schema extractions. The gateway always charged this bucket but never returned it, so the published per-source credit fields could not add up to credits_used.

Changed

  • Internal GET /v1/bulk is cursor-paginated and its 100-job cap is gone (#4508). The generated listBulk types record a changed response shape: the operation used to return a bare array of at most the 100 newest jobs, and now returns the same page envelope GET /v1/monitors does, { data, next_cursor, has_more }, plus limit and cursor query parameters (limit is 1 to 100, default 25). A status filter and a cursor combine in one request: status narrows the set and the cursor addresses a position inside it. The operation's summary also no longer says "in-flight", which was wrong — it has always returned jobs of every status, cancelled included. This is a breaking change to the wire shape of an x-sdk-scope: internal operation; no SDK method calls it, so no published method signature moves.
  • GET /v1/logs billing_multiplier is the bandwidth factor, not a credit count (#4469). The generated logs-row type records the changed meaning of a public field: the value is now billable_bytes / raw_bytes rounded to 2 decimal places — the factor the request settled at, carrying the endpoint, cache-hit and clean-markdown factors, and the flat 1.5 Phase-1 factor on screenshots and recordings. It used to be the row's persisted credits_charged, an account credit count charged per capture option, so two screenshots billed at the same 1.5 factor reported 2 and 1. The 2 decimal places are load-bearing rather than cosmetic: billable_bytes is itself rounded, so an odd byte basis would otherwise surface as 1.500001943158721 instead of 1.5 on roughly half of all such rows. Every factor the field can report is a multiple of 0.25, so 2 places is exact for any basis of 100 bytes or more, and raw_bytes × billing_multiplier rounded to the nearest byte still reproduces billable_bytes. The billable_bytes doc comment therefore no longer says it is NOT raw_bytes × billing_multiplier. The field is also null on more rows than before: a search or webhook row has no byte basis, a capture whose Phase 1 was skipped has a zero one, and a capture old enough to predate the Phase-1 basis being recorded has none at all. One residual is documented rather than fixed: a markdown /v1/fetch row completed before the billing decomposition was recorded reports 1.0 where settlement charged 1.5, because nothing on such a row distinguishes it. The type shapes are unchanged (number | null), and no SDK method calls this endpoint.
  • An all-null /v1/schema extraction your schema allows is validated again (#4490). The generated /v1/schema 200 response type records the exception to #4455's rule: an output whose every leaf is null still arrives validated: true with the object intact, and is charged, when the schema you sent declares "null" in the type union of every leaf the output returned as null. Only that form counts — nullable: true, an enum containing null and const: null do not, and a response that is itself null rather than an object is never accepted however the root is declared. Everything else stays the #4455 miss. The data doc comment now also states the other half: a validated: true response always carries an object, never null. The type shapes are unchanged; only their doc comments moved.
  • An all-null /v1/schema extraction is reported, not validated (#4455). The generated /v1/schema 200 response type documents a new additive warning code, empty_extraction: a schema-valid output whose every leaf is null now arrives with data: null and validated: false, and is not charged the per-request credits. Previously it arrived as validated: true with the all-null object and no warning. The data and validated doc comments now also say that every validated: false response carries data: null; the old comment claimed exhausted retries returned the last raw output, which the API never did. The type shapes are unchanged; only their doc comments moved.
  • An inline POST /v1/scrape failure is documented with a persisted id (#4454). The gateway now records a sync scrape that ends upstream_error or blocked as a failed scrape job, so its failed envelope carries that job's sc_ id instead of null, scrape.get(id) returns the same envelope, and the request appears in the dashboard Logs. The generated createScrape 200 and ScrapeErrorEnvelope doc comments and the scrape.create doc comment no longer say id: null. The type shapes are unchanged (id stays string | null), and scrape.create still throws ByteKitError on that envelope exactly as before.
  • Internal GET /v1/logs types list search and webhook rows (#4460). Both endpoint enums add search and webhook; a row's source is now string | null. No SDK method calls it.

[0.10.0] - 2026-09-11

Changed

  • A terminal POST /v1/scrape failure now throws instead of resolving (#1870). The gateway used to report an upstream fetch failure — unreachable host, refused connection, or bot protection that defeated every proxy tier — as HTTP 503; it now answers HTTP 200 with the canonical failed envelope (id: null), because a definitive upstream outcome is a result, not a transport error. !response.ok was this SDK's only failure seam, so without this change that body would resolve like a success and silence every caller's catch. scrape.create (and any request / requestWithResponse call to POST /v1/scrape) now raises ByteKitError carrying the envelope's code and message, the actual status 200, the response headers, and error.details when one is supplied. A legacy 503 from a not-yet-upgraded gateway still throws through the unchanged non-2xx decoder, so the SDK works against both. Upgrade note: if you inspected result.status === 'failed' after scrape.create, move that handling into a catch.
  • Everything else is untouched: scrape.get(id), bulk item results, and every other endpoint still deliver failed as data; requestRaw (/v1/fetch) never parses JSON and is unaffected; 204 and empty-body handling are unchanged. request() now delegates to requestWithResponse() so the guard has a single application site — identical in every observable respect but the return value, as both were already documented.

[0.9.4] - 2026-09-06 [NEVER PUBLISHED]

This version was cut but NEVER PUBLISHED — npm latest goes from 0.9.3 straight to 0.10.0. Publication is triggered by CI on main from the version the manifest declares, and 0.10.0 (#1870) superseded 0.9.4 in the tree before any release reached main. Everything below reaches consumers in the 0.10.0 artifact, which contains all of it; nothing is installable at 0.9.4: npm i @hunt-labs/[email protected] fails.

Added

  • Bulk fetch, scrape, and mixed-result reads now accept optional cursor, limit, and status page parameters while preserving the existing (id, requestOptions?) call shape (#3717). Generated response types publish stable item identities, page cursors and completeness, whole-job counters, exact mixed artifact keys, artifact warnings, bounded webhook recovery metadata, and the artifact_unavailable error contract.

Fixed

  • Corrected the artifact-URL fallback documentation on GET /v1/bulk/{id}/screenshots (image_url, content_url, markdown_url) and GET /v1/logs (artifact_url) (#3717). Those fields were documented as null-plus-artifact_unavailable when no signature can be minted; the servers have always returned the row's stored value unchanged instead — null while nothing is stored, and otherwise the private unsigned storage URL, which answers 401 — and raise no warning. GET /v1/fetch/bulk/{id} does null the field and report artifact_unavailable, and its documentation is unchanged. No server behaviour changed and no type moved; the descriptions now match what the endpoints return.

Changed

  • Regenerated src/generated/api.d.ts against the corrected topup_balance_bytes description on GET /v1/usage (#4251). The description read "Remaining purchased top-up bytes", which contradicted its own "Equal to topup_bytes" clause: the value covers purchased top-up lots and the one-time starter grant, which is what the endpoint has always returned. The clause recording that the figure is already net of consumed overage, and the "Present only in dual/bytes billing mode" clause, are both unchanged and still true. Types are unchanged and no server value changed; only the documentation on them moved.
  • Regenerated src/generated/api.d.ts against the clarified cache_age_s description on ScrapeSuccessEnvelope (#4237). The description now records that GET /v1/scrape/{id} reports cache: hit without cache_age_s: the age is a property of the replay that happened at completion time, is not persisted on the row, and cannot be derived honestly afterwards. The one-directional "Present only on cache: hit" clause is unchanged and still true. Types are unchanged; only the documentation on them moved.
  • Regenerated src/generated/api.d.ts against the completed cache presence list on ScrapeSuccessEnvelope (#4203). The field's description now names GET /v1/scrape/{id} as a third delivery path, alongside the synchronous response and webhook delivery: a polled scrape now reports the cache disposition its row recorded, and omits the field — rather than defaulting it to miss — for a row that recorded none. Types are unchanged; only the documentation on them moved.
  • Regenerated src/generated/api.d.ts against the corrected PUT /v1/billing/auto-topup request body (#3984). monthly_cap_cents is now typed number | null: an explicit null is the wire representation of "no monthly cap", and it is what the endpoint has always accepted — the request schema simply never documented it, while the 200 response schema already did. Omitting the key while enabled is true is still rejected with 422, so send null rather than dropping the field. No server behaviour changes; the type widens to match what the API already accepted.
  • Regenerated src/generated/api.d.ts against the corrected monitor scrape_options documentation (#4024). The description no longer lists headers among the accepted fields — POST /v1/monitors never honoured a caller-supplied headers option and now rejects it with the same 400 every other capture endpoint returns for the field. Types are unchanged; only the documentation on them moved.
  • Regenerated src/generated/api.d.ts against the completed X-Scrape-Proxy-Tier-Label omission list (#4030). The header's description now names a third case in which no tier is reported: a response replayed from stored bytes after this request's own cache lookup missed, which arrives as X-Scrape-Cache: miss and likewise ran no proxy tier. The two cases already documented — a cache hit, and responses the fast path did not serve — are unchanged. Types are unchanged; only the documentation on them moved.
  • Regenerated src/generated/api.d.ts against the qualified X-Scrape-Proxy-Tier-Label omission list (#4132). The third case that list records — a response replayed from stored bytes after this request's own cache lookup missed — is not reachable through this API: every caller of the internal fetch path closes the handler branch that produces it, and the branch survives only as a deploy-skew safeguard. The description now says so rather than presenting the case as ordinary behaviour. The two reachable cases — a cache hit, and responses the fast path did not serve — are unchanged. Types are unchanged; only the documentation on them moved.

[0.9.3] - 2026-09-04

Changed

  • Regenerated src/generated/api.d.ts against the corrected queued-scrape status_url documentation (#3925). status_url is a root-relative path, never an absolute URI: join it onto the API base URL you called (https://api.bytekit.com, no trailing slash) to get the URL to poll. The spec had declared it format: uri with an absolute example, which no server has ever emitted; the emitted value is unchanged, so no client behaviour moves.
  • Regenerated src/generated/api.d.ts against the documented proxy-tier response headers (#3926). X-Scrape-Proxy-Tier-Label now documents that it reports the tier that ACTUALLY served a scrape — after any escalation, never the tier first attempted — that it is diagnostic only, and that it is omitted on a cache hit, which replayed stored bytes and ran no tier. The deprecated numeric X-Scrape-Proxy-Tier documents the legacy 1 datacenter / 2 residential / 3 mobile vocabulary and its omission for the lean-browser tier. Types are unchanged; only the documentation on them moved.

[0.9.2] - 2026-09-03

Changed

  • Breaking (API): the scrape formats vocabulary is now raw | markdown | links | images (#3775). The unprocessed-source format is called raw — it has returned JSON, XML and plain text alongside HTML since #1853, so the old raw-HTML name misdescribed it — and the cleaned-article-HTML format is gone. The server accepts no alias for either retired value, so a request naming one now comes back 422 with the new set in the message. Update formats and read the content back from formats.raw:

    const result = await client.scrape.create({ url, formats: ['raw', 'markdown'] });
    console.log(result.formats.raw);
  • Regenerated src/generated/api.d.ts against the documented 402 contract (#3784). Every quota-enforced operation now declares 402 through one shared PaymentRequired response, and the X-Quota-* headers are referenced by the operations that actually emit them. Types only — no runtime behavior changed.

  • The generated events field description on the scrape request type now documents the widened server default (queued, completed, failed) and that a caller may send an explicit subset to receive fewer deliveries (#3777). Types/docs only — this regeneration followed a server-side default change; the SDK carried no logic tied to the old default.

[0.9.1] - 2026-08-25

Fixed

  • The generated paths type for the 200 of GET /v1/scrape/bulk/{id} now declares items (#3632). The route builds one body and varies only the status code, so the terminal 200 has always carried the per-URL scrape envelopes — but only the 202 documented them, so scrape.bulk.get() typed the results away at exactly the poll that has them. Types only: the field was already on the wire, and server behavior is unchanged.

  • The generated paths types now describe the 202 that GET /v1/scrape/bulk/{id}, GET /v1/fetch/bulk/{id} and GET /v1/bulk/{id} already serve while a bulk job is non-terminal (#3533). Each of the three returns 202 with the job envelope until the job is terminal and 200 afterwards, but only the 200 was declared, so anything typed off the spec had no shape for the response the very first poll receives. Types only: scrape.bulk.get(), fetch.bulk.get() and bulk.get() already resolved the body on any 2xx, and server behavior is unchanged.

Changed

  • The generated topup_balance_bytes description on GET /v1/usage now states that the value is the remaining purchased top-up bytes, already net of any overage consumed against them (#3459). Types only: the field's name, type and wire format are unchanged. The server previously subtracted the consumed overage a second time when computing this value, so the number a client reads is larger than before for accounts that ran past their included allowance.

[0.9.0] - 2026-08-20

Removed

  • BREAKING: the cookies and headers request-body fields are gone from every capture endpoint (#3242) — /v1/scrape, /v1/scrape/bulk, /v1/screenshots, /v1/recordings and /v1/schema. A request that carries either field is rejected with 400 validation_error; it is not ignored, and there is no deprecation window, feature flag, or compatibility header. ScrapeOpts.cookies, ScrapeOpts.headers and the exported ScrapeCookie type are removed from the SDK surface, so a call that still passes them no longer compiles.

    // before — compiled, and the values were forwarded to the target site
    await client.scrape.create({ url, cookies: [{ name: 'session', value: 'abc' }] });
    // after — does not compile; the same body over raw HTTP is a 400 validation_error
    await client.scrape.create({ url });

    ByteKit continues to manage its own cookie jar and request headers internally; what was removed is the caller's ability to inject or override them.

Changed

  • The generated types now declare the 503 response of GET/POST /v1/fetch (#3372). /v1/fetch answers 503 on two conditions and both carry the ordinary Error envelope, so a typed consumer reading operations['getFetch']['responses'] / ['postFetch'] can handle it instead of meeting an undeclared status at runtime. Read error.code to tell the two apart, because they disagree about error.http_status:

    • upstream_error — the target returned HTTP 502, relayed at 503 because Cloudflare replaces a clean origin 502 with its own error page. error.http_status stays 502, the true upstream status, and is the only surviving record of what the target answered.
    • internal_error (Fetch service unavailable.) — the fetch service was unreachable. There is no upstream, so error.http_status reports 503, matching what was served.

    No runtime behaviour changes, and every other documented /v1/fetch status is unaffected.

[0.8.0] - 2026-08-19

Removed

  • BREAKING: billing_multiplier and billed_bytes are gone from the scrape success envelope (#3143). Read billing instead — a BillingNode that names every factor applied to the charge, the multiplier they compose to, and the resulting billed bytes:

    // before
    const spend = res.billed_bytes; // flat, base-endpoint only
    // after
    const spend = res.billing?.billed_bytes;
    const why = res.billing?.factors; // [{ name: 'endpoint', value: 1.5, reason: 'scrape_md' }, …]

    This is a correction, not a rename: the two surfaces reported different numbers for the same charge. The flat pair carried the BASE-endpoint composition, so a markdown cache hit read 0.5 / 400 while the real charge — and the X-Billing-Multiplier header — was 0.75 / 600; an html-only miss read the per-format credits estimate 1.5 / 150 against an endpoint factor of 1.0 / 100. The flat pair also never carried the clean_markdown ×3 surcharge, so on a successful clean it under-reported the amount billed by two thirds. billing.multiplier composes every applied factor behind exactly one rounding, so it is the number that predicts your invoice.

    billing is optional and is absent on resources that completed before the billing node existed. A pre-node row omits the field rather than reporting a zero node, because a zero node would assert that a real past charge was zero. Test the node for presence (res.billing !== undefined, or equivalently a plain truthiness check — billing is an object or undefined, and an object is always truthy, so the two are the same test here). The hazard is a condition that reaches into the node: if (res.billing?.billed_bytes) silently skips a genuine zero-byte charge, because 0 is falsy. Check that the node is present, then read billed_bytes from it.

    // wrong — a real 0-byte charge is falsy and disappears
    if (res.billing?.billed_bytes) record(res.billing.billed_bytes);
    // right — presence first, value second
    if (res.billing) record(res.billing.billed_bytes); // 0 is a charge, not an absence

    Two endpoints deliberately carry no node: /v1/search is billed per credit and reports credits_used, and a single /v1/fetch reports its charge through response headers only. The /v1/logs row field named billing_multiplier is a different field and is not the envelope field this release removed. It carries the BASE settled credits_charged factor, not the composed real-endpoint value the X-Billing-Multiplier header reports — a markdown cache hit reads 0.5 on the logs row and 0.75 on the header.

Changed

  • Generated ScrapeErrorEnvelope types now describe terminal asynchronous scrape.failed webhook delivery (#1856). No runtime SDK behavior or API shape changed.

Fixed

  • GET /v1/logs entries now carry billing_multiplier, as a JSON number (#3022). The gateway had been serializing that value under an undeclared key credits, and as the raw Postgres NUMERIC(10,2) string ("1.50"), while the OpenAPI spec — and therefore the generated types in src/generated/api.d.ts — declared billing_multiplier?: number | null. Anything typed against this package read undefined for the field.

    Breaking at the wire level: the undeclared credits key is gone from every /v1/logs entry. Code that reached past the generated types to read credits must read billing_multiplier. No SDK method exposes this endpoint, so the only change inside this package is the regenerated type description.

  • change_threshold (on monitors.create/monitors.list/monitors.get/ monitors.update responses) and change_pct (on monitors.captures.list entries) now serialize as JSON numbers, not strings (#3021). Both are Postgres NUMERIC(5,2) columns; the gateway passed the raw driver string through unconverted ("5.00" / "7.25") while the OpenAPI spec's request schemas already declared number — only the response schemas had been (incorrectly) edited to match the buggy string output. This release fixes the gateway serializer and corrects the two response schemas to number / number | null, so spec and runtime now agree.

    Breaking at the wire level: client.monitors.create(), client.monitors.list(), client.monitors.get(), and client.monitors.update() now return a JS number for change_threshold instead of a string (screenshot-type monitors only); the create response's MonitorCreateResponse and every other monitor route serialize through the same MonitorResponse schema, so all four are affected identically. client.monitors.captures.list() now returns a JS number for change_pct instead of a string. Code doing parseFloat(monitor.change_threshold) or template-literal interpolation keeps working; strict typeof x === 'string' checks or string methods (.trim(), .padStart()) will now throw. change_pct is null (not 0) when there is no prior capture to diff against, unchanged from before. No SDK method-level shape changed otherwise — the only change inside this package is the regenerated type description.

[0.7.4] - 2026-08-19

Changed

  • gen:version no longer runs as part of build (#3144). scripts/gen-version.ts writes src/version.ts, a tracked source file, and it was chained into the build script. That made a build silently rewrite committed source whenever the file disagreed with the manifest: local runs went green on a tree CI could fail, the resulting M src/version.ts was indistinguishable from routine build noise, and in CI the outcome depended on whether Turbo's build cache hit (restoring dist/** without re-running the generator) or missed. build now starts at rm -rf dist ..., so building this package touches nothing tracked. This changes prepack / prepublishOnly behavior too: both call build, so the publish path now reads the committed src/version.ts rather than regenerating it, and the committed value is what ships.
  • New gen:version-check script, pnpm gen:version && git diff --exit-code src/version.ts, mirroring the existing gen:types-check in the same manifest. It runs in .woodpecker/ci.yaml's types-drift step, so a manifest bump without the matching src/version.ts regeneration now fails a named PR check instead of reaching staging. src/__tests__/version-parity.test.ts stays in place as the in-suite assertion; the gate is what makes its failure reachable before merge rather than after.

No runtime or type changes. SDK_VERSION and the X-ByteKit-Client: sdk_ts/<version> affix behave exactly as before.

[0.7.3] - 2026-08-18

Changed

  • XFetchContentLength's description no longer calls the header a wire-byte count (#3133). X-Fetch-Content-Length is the byte length of the response body actually served: on a passthrough /v1/fetch (no format) that is the relayed upstream payload, so it does equal the upstream wire bytes — but with format=html or format=markdown the gateway overwrites it with the transformed document's length, which is typically far smaller than what is billed. The old description therefore read as a billing discrepancy on every transform request. Billing is unchanged and has always used the upstream compressed wire bytes reported by X-Raw-Bytes / X-Billable-Bytes. Generated-type comments only; no type or runtime behavior changes.

[0.7.2] - 2026-08-17

Changed

  • ScrapeResponses.202's headers no longer advertise X-Raw-Bytes, X-Billable-Bytes, or X-Billing-Multiplier (#3068). The gateway's async /v1/scrape path only ever emitted X-Scrape-ID and X-Scrape-Status: queued on a 202 — a queued scrape has no finalized byte charge yet — but docs/api/openapi.yaml documented the billing-transparency trio there too, so src/generated/api.d.ts typed three response headers the runtime never sends. This corrects the generated type to match the gateway's actual behavior; it does not change any runtime behavior. The synchronous 200 response, and /v1/screenshots' 202 (which does finalize a charge before responding), are unaffected.

[0.7.1] - 2026-08-11

Changed

  • Generated types follow the spec's corrected error and format vocabularies (#2959). Two vocabulary corrections in docs/api/openapi.yaml flow through src/generated/api.d.ts:
    • ScrapeErrorEnvelope.error.code no longer lists quota_exhausted. The API has emitted quota_exceeded since #1539 and there is no throw site for quota_exhausted anywhere in the gateway, so this removes a code that has never been sent — not one that stops being sent. A consumer switching on it was matching a branch that could not be reached; the live code for an exceeded quota is quota_exceeded. invalid_url, which the API does emit, is unchanged and still listed.
    • createBulk's formats now spells the raw-HTML value raw_html, matching ScrapeRequest and createScrapeBulk. /v1/bulk previously accepted only the internal rawHtml spelling and the spec documented that; it now accepts both, so the one endpoint that disagreed with the rest of the surface no longer does. rawHtml keeps working.

Added

  • ByteKitError.headers — the failing response's headers, so you can write your own backoff (#2957). This SDK performs exactly one fetch per call and never retries, deliberately, so that retry policy stays the caller's. But the policy was unwritable: ByteKitError carried no headers and requestWithResponse returns headers only on 2xx — which a rate-limited request never reaches — so Retry-After and the X-RateLimit-* family were unreachable through the SDK at every entry point. err.headers closes that:

    if (err instanceof ByteKitError && err.status === 429) {
      const waitMs = Number(err.headers?.get('Retry-After') ?? 1) * 1000;
    }

    It is the runtime's own Headers object rather than a snapshot, so get() stays case-insensitive whatever casing the wire used. Strictly additive: it is optional and last on the constructor, follows the same declare discipline details uses, and an error constructed without one carries no headers key at all — 'headers' in err, spreads and JSON.stringify are unchanged for every pre-existing construction, as are status, code, message, details and both 2xx return shapes. Only response headers are exposed; Authorization rides on the request and is never reachable through it. ByteKitConnectionError is unchanged — no response, no headers to carry.

Fixed

  • screenshots.getWithResponse's docstring no longer promises headers that endpoint never emits (#2957). The JSDoc — the one an IDE surfaces on hover, and the one the published SDK reference is generated from — claimed the documented X-Screenshot-* / X-Raw-Bytes / X-Billable-Bytes / X-Billing-Multiplier headers "are on headers". They are not: GET /v1/screenshots/{id} is declared with no X-* response headers, which is what ByteKit.requestWithResponse's own doc comment and the README have said all along, and what a live GET confirms. Documentation only — no signature, return type or runtime behavior changed. screenshots.createWithResponse (POST /v1/screenshots) really does emit that family, and its docstring still says so.

  • README documents the signal-vs-timeoutMs precedence and the requestTimeoutMs escape hatch (#2957). Passing your own AbortSignal silently disables the constructor's timeoutMs — the signal becomes the request's only abort source, so a client built with timeoutMs: 200 waits indefinitely if the signal never fires. That precedence was documented on RequestOptions.signal in the types, where only a reader already inspecting the option would meet it, and nowhere in the README. A new "Timeouts and cancellation" section states it with the full four-row precedence table and shows requestTimeoutMs composed alongside a caller signal (the request then aborts on whichever fires first). No behavior changed — the precedence is what it always was.

  • ScrapeWarning.code now carries all 12 codes the server can emit (#2948). The generated union in src/generated/api.d.ts listed 10. low_quality_extraction and fragment_menu_stripped — added to the server-side emitter union by #1701 — never reached docs/api/openapi.yaml, so every regeneration since faithfully reproduced the gap and the exported type was non-exhaustive: a switch over warning.code that TypeScript proved exhaustive would silently fall through on a warning the API really returns. Widening the spec and regenerating fixes the type; no runtime code changed. The other 10 codes are unaffected.

  • Gateway wire-type correction: billing_multiplier is a JSON number on scrape.get() again (#2954). GET /v1/scrape/{id} (and GET /v1/scrape/bulk/{id} items) were serializing billing_multiplier as a JSON string (e.g. "0.50") instead of a number — the pg driver returns the underlying NUMERIC(10,2) column as a string, and the gateway's GET serializer passed it through unconverted, while POST /v1/scrape's create-path value (computed in JS) was already correct. Fixed upstream in the gateway; this SDK required no code change, only removal of the KNOWN DRIFT (#2426) JSDoc on ScrapeResource.get() that documented the workaround — the static number type this SDK has shipped all along is accurate again. A caller that was defensively coercing with Number(...) per the old note does not need to change anything; a caller relying on the exact string form "0.50" (rather than doing arithmetic) will now see 0.5.

Changed

  • Alignment release — no consumer-visible change in the alignment itself (the #2948 type widening above is a separate, deliberate change under this same version). 0.7.0 published to npm at 2026-08-11T21:17:42.790Z. Two doc comments, in src/client.ts and src/version.ts, were reworded after the commit that introduced the 0.7.0 version string — they had re-spelled the X-ByteKit-Client header name in prose, which issue #2905's pre-merge gate 1 counts. No exported function, type, or runtime behavior differs from 0.7.0.

    This release exists because check-release-drift (#2680) is version-gated and anchors on the commit that introduced a version, not on the commit whose tree was actually published. Once 0.7.0 reached npm, any shipped-file change after that anchor reads as post-release drift. Bumping moves the anchor past those comments. The published 0.7.0 artifact is not stale in any way a consumer can observe — the divergence is comment text only.

  • Corrected a false legacy-name claim carried since the rename (#2950). The 0.7.0-era note below said the prior rapidcrawl/@rapidcrawl/sdk name "remains installable at its last version but is deprecated." Verified 2026-08-13: both rapidcrawl and @rapidcrawl/sdk return npm 404 ("is not in this registry"), and npm deprecate against either fails with the same 404 — there is no published version left for a deprecation notice to attach to, so the claim never held. No registry stub or deprecation was possible to publish; root CLAUDE.md carried the same false claim and is corrected alongside this entry. Users of the old names should install @hunt-labs/bytekit-sdk directly. No exported function, type, or runtime behavior changes.

[0.7.0] - 2026-08-11

Added

  • Every outbound request now carries X-ByteKit-Client: sdk_ts/<version>. The SDK identifies itself with the sanctioned client_surface marker (issue #2905, slice 2 of #2701), so requests made through @hunt-labs/bytekit-sdk are distinguishable from raw API calls in ByteKit's own analytics instead of all resolving to api_direct. sdk_ts is a member of the closed client_surface vocabulary declared in issue #2904; the version affix is this package's own version, baked in at build time from package.json.

    Set at exactly ONE place — the header literal inside the private transport every request funnels through — so it rides on the JSON path, the raw /v1/fetch path and every resource method alike. Nothing else about a request changed: same Authorization and Content-Type, same body bytes, same client-side abort budget, same error taxonomy, same return shapes. The marker carries no caller data and cannot be set, overridden or read by callers — there is no headers option on ByteKitOptions or RequestOptions.

    Minor rather than patch because the wire format of every request changes, even though no documented API does. A gateway that does not yet recognize the header ignores it.

  • SDK_VERSION — an internal, build-time-generated constant (src/version.ts, written by pnpm gen:version from package.json and gated by src/__tests__/version-parity.test.ts). It backs the marker's version affix. Not part of the public export surface; package.json's version field remains the only version consumers should read.

Unchanged

  • Still ZERO runtime dependencies. The client_surface vocabulary lives in an internal, unpublished workspace package, so the emitted value is a string literal here and the binding to that registry is a test-only import. dependencies stays empty, asserted by src/__tests__/package-manifest.test.ts.

[0.6.2] - 2026-08-11 [NEVER PUBLISHED]

This version was cut but NEVER PUBLISHED — npm latest went from 0.6.1 straight to 0.7.0. The two documentation/regeneration changes below reached consumers in the 0.7.0 artifact, which contains all of them. Nothing under this heading is missing from 0.7.0, and nothing is installable at 0.6.2: npm i @hunt-labs/[email protected] fails.

Why it was stranded: 0.6.2 was cut for two src/generated/api.d.ts regenerations, and the next change (issue #2905's new outbound request header) altered the wire format of every request — additive behavior that belongs at a MINOR, not a patch. Folding it into the still-unpublished 0.6.2 would have shipped a new request header to consumers pinning ^0.6.1 unannounced, so 0.7.0 was cut instead and this heading annotated per ci/scripts/check-release-drift.sh's reconciliation form (#2680).

Documentation

  • AccountSessionResponse.has_activity now documents what it actually means. The generated type's JSDoc previously said the flag was true once the account had "ever recorded billable activity", which read as "ever made a request". It is true only once a request has successfully completed and been charged credits — a request that failed does not flip it, even when its bandwidth was billed (issue #2890). The doc comment also now states the all-time (never windowed) scope and the up-to-one-rollup lag, so a caller polling it after a first request knows false can mean "not visible yet" rather than "never made a request". Regenerated from docs/api/openapi.yaml; the emitted TYPE is unchanged (has_activity: boolean), so this release is source-compatible in both directions.

Changed

  • Regenerated src/generated/api.d.ts from docs/api/openapi.yaml after issue #2904 declared the optional X-ByteKit-Client request header as components.parameters.ClientSurfaceMarker. The regeneration is ADDITIVE and touches no method signature: one entry appears under components["parameters"], because the parameter is deliberately $ref-ed from no operation. No consumer-visible change — no exported function, type or runtime behavior differs from 0.6.1. The entry exists because check-release-drift (#2680) requires a version bump for any post-release change to a shipped file, and the publish pipeline is version-gated.

    Documented under 0.6.2 alongside the #2890 regeneration above, rather than under a 0.6.3 of its own, because 0.6.2 has been CUT BUT NOT PUBLISHED (npm latest is 0.6.1). Both regenerations therefore ship in the same single artifact. Numbering this 0.6.3 would leave 0.6.2 documented and unpublishable — the version-gated pipeline would publish 0.6.3 and skip it — which is precisely the phantom-version class #2680 exists to kill, and check-release-drift rejects it by name.

[0.6.1] - 2026-08-10

Changed

  • Version-alignment release after the 2026-08-10 publish rescue: the published 0.6.0 artifact was built from the main promotion HEAD, which already contained every change documented under the 0.6.0 heading — but that commit postdates the one that introduced the 0.6.0 version string, so the check-release-drift gate (#2680) correctly reported shipped drift on every subsequent CI run (#2829). 0.6.1 realigns manifest ↔ registry ↔ version-introducing commit. No consumer-visible change relative to the published 0.6.0.

[0.6.0] - 2026-08-07

Added

  • ByteKitError.details — the API error envelope's structured rejection reason is now reachable from the thrown error. A 422 validation_error carries { formErrors, fieldErrors }, so a caller can finally read WHICH request field was rejected rather than only that the request was invalid; a 400 invalid_url carries a flat { url, field, reason }. Typed as ByteKitErrorDetails | undefined where ByteKitErrorDetails = Record<string, unknown> (also exported), deliberately loose because the contents vary by error code — narrow the member you read. Strictly ADDITIVE: details is an optional FOURTH constructor argument, so existing three-argument new ByteKitError(status, code, message) construction is unchanged, and status / code / message are byte-identical. Populated only from the canonical { error: { … } } envelope — a flat vendor-masked body (/v1/search 502 {"error":"search_provider_error"}) carries no structured detail and grows none. http_status and failed_at are deliberately NOT exposed: http_status would duplicate the existing status property, and a duplicated field that can disagree with its twin is a bug surface. (#2685)

  • A tsc --noEmit-checked parity pin in src/types.ts ties the hand-authored search types (SearchOpts, SearchResult, SearchResultImage, SearchResponse — the SDK's only hand-authored request/response types) to the generated operations['createSearch'] request and 200 response. It asserts key-set and optionality parity in both directions and assignability in each meaningful direction, so a future spec change that search's types do not follow fails the build instead of shipping. Value types are erased before the parity comparison, so a deliberate hand-authored refinement (a branded id, a narrowed union) cannot trip it. Type-level only — the assertions erase at emit and add nothing to dist/. (#2678)

Fixed

  • README: the fetch sample's res.cache type is corrected to include null. The documented type read "hit" | "miss" | "bypass", but FetchResult.cache is FetchCacheStatus | null (null when the response carried no recognizable X-Fetch-Cache header) — the one strict-tsc error a real consumption file copied from the README ever produced. The line is now a TYPED ASSIGNMENT rather than a comment, so the README compile guard (src/__tests__/readme-compiles.test.ts) typechecks it: comment-vs-type drift on this field now fails the build instead of shipping. Documentation only — no runtime or type change. (#2685)

  • CHANGELOG: the [0.1.0] heading's malformed - 2026 date is now - 2026-04-24, the commit date of caf17e56b, which introduced packages/sdk-typescript/package.json at version 0.1.0. Every other heading in this file already carried a full ISO date; this one broke the Keep a Changelog ## [x.y.z] - YYYY-MM-DD form. (#2685)

  • Test-only: nine SDK test files mocked wire bodies the API can never send. A monitors.create 200 (the operation declares 201 as its only success arm), a scrape 200 carrying data: { markdown } (the envelope has no data member — content lives under formats), a screenshots.getWithResponse mock asserting X-Screenshot-* headers that GET /v1/screenshots/{id} declares on no arm, error bodies omitting the Error schema's required status: 'failed' / http_status / failed_at, and twelve 402/429 arms on operations that declare neither status. Each is rewritten to a spec-valid shape with a comment naming the arm it now matches, and every test keeps its original assertion target. No shipped code changed. (#2685)

  • SearchResult.snippet is now optional (snippet?: string) — BREAKING AT COMPILE TIME. type: 'images' search results omit snippet entirely (the spec has always declared a result's required set as [position, title, url]), but the SDK typed it as a required string. hit.snippet.length therefore compiled under strict TypeScript and threw TypeError: Cannot read properties of undefined on every images hit — the TypeScript twin of the Python SDK's #2579 images KeyError. Consumers reading .snippet unguarded now see string | undefined and must narrow it (if (typeof hit.snippet === 'string') or hit.snippet ?? ''), exactly as they already do for date. No runtime behavior and no wire bytes change: the field was already absent from those payloads. web/news results are unaffected — they still carry snippet, and a narrowed read yields the same string. (#2678)

  • Path parameters are now percent-encoded. Every id-bearing method built its request path with a raw template literal, so a crafted or interpolated id silently retargeted the request: scrape.get('sc_/../../account') addressed GET /v1/account once WHATWG URL parsing collapsed the ../ segments, and monitors.get('mon_x?admin=1#frag') turned the id into a real query string plus fragment. The branded id types constrain only the PREFIX of an id, never its tail, so they were never a defence. An id now travels as one percent-encoded path segment — a malformed one 404s server-side instead of reaching a different resource. Ports the Python SDK's #2593 fix. (#2677)

Changed

  • Request paths are assembled by a single internal buildPath() helper in client.ts rather than by per-resource template literals. It encodes SEGMENTS only — every fixed segment and every / separator stays literal — and encodes each value exactly once, so an id that already contains a % is not double-encoded. A well-formed id (sc_<hex>, mon_<hex>, …) produces a byte-identical URL to before, on every id-bearing method: encoding is a no-op for those characters, so no existing call changes on the wire. (#2677)

  • Four of the guards shipped above could not fail on the change they appeared to protect, and now can. buildPath()'s two TypeError throws — the fail-loud that replaced the compile-time checking a template literal used to give — each survived its own removal with the whole suite green, and are now pinned by tests separated by input shape rather than by message text. The in-suite path-helper gate read its directory non-recursively while the shell grep it reproduces is -r, so an offender one directory down was invisible to it; it now walks the whole of src/, a strict superset of the shell gate's set. The images search-result literal's key ABSENCE (no snippet, no date) was held only by a one-directional subtype check, which an extra key cannot break; it is now pinned by explicit key-absence assertions plus an exact key-set assertion. And the spec's error.details — read at one site in client.ts off a hand-written inline type — is now tied to the generated schema by a PRESENCE pin, so dropping the field from docs/api/openapi.yaml fails the build. Every pin is key-level, never value-level, so a deliberate hand-authored value refinement still cannot trip it. Test- and type-level only: no runtime behavior, no exported type, and nothing in dist/ changes. (#2793)

[0.5.0] - 2026-07-29

Added

  • src/generated/api.d.ts now models the 400 responses that the gateway's new request-body-size guards can return. operations['stripeWebhook']['responses'] gains a 400 with this route's bespoke { error?: string } body, and operations['retryWebhookDelivery']['responses'] gains a 400 carrying the unified components['schemas']['Error'] envelope. Both are purely additive — no existing member changed shape — so every current call site still compiles. (#2640)

Changed

  • Regenerated src/generated/api.d.ts against the merged spec. Beyond the two added 400 members above, the delta is JSDoc only: every guarded endpoint's description now carries a **Limits:** block stating its byte ceiling and the request_too_large error code (#2640), and POST /v1/bulk's description documents defaults.type inheritance and the 422 on unrecognized keys (#2643). No runtime code changed. (#2640, #2643)

[0.4.0] - 2026-07-24

Added

  • RequestOptions (per-call requestTimeoutMs + caller signal) is now accepted by every resource method, not just scrape.create / scrape.get. It is always a trailing optional parameter — on screenshots.create it is the third positional, after params — so every existing call site compiles and serializes byte-identically. Per-call options continue to beat the constructor's timeoutMs default; omitting them leaves the constructor default in force exactly as before. (#2547)
  • ApiResponse<T> ({ data, headers, status }), a generic client.requestWithResponse<T>() transport, and three typed siblings — sitemap.getWithResponse, screenshots.createWithResponse, screenshots.getWithResponse — make the documented X-Raw-Bytes, X-Billable-Bytes, X-Billing-Multiplier and X-Screenshot-* response headers reachable. This surface is purely additive: it issues the identical request with the identical options as the base method, and no existing return type changes. (#2547)
  • MonitorStatus and ScrapeCookie are exported from the package entrypoint. (#2547)

Changed (BREAKING at compile time)

  • ScrapeOpts.cookies is now ScrapeCookie[], derived from the spec's Cookie schema (name and value required; domain, path, secure, httpOnly optional), instead of an open bag of untyped records. A cookie missing name used to compile and fail as a 422 on the wire; it is a compile error now. Serialization is unchanged. (#2547)
  • ListMonitorsParams.status is now the spec enum ('active' | 'processing' | 'paused' | 'cancelled' | 'suspended'), derived from the generated listMonitors query type, instead of string. A typo like 'activee' used to compile and 422; it is a compile error now. The serialized query string is unchanged. (#2547)

Fixed

  • timeoutMs / requestTimeoutMs (and a caller-supplied signal) now cover the response body read, not just the time to first byte of headers. The abort budget was armed for fetch and cleared the moment headers arrived, after which request()'s response.text(), requestRaw()'s response.arrayBuffer(), and the error-body response.json() on a non-2xx all read the body unbounded — so a server that answered with headers and then stalled its body hung the caller forever. A mid-body abort now surfaces as ByteKitConnectionError with code: 'timeout' (budget) or 'aborted' (caller signal) instead of a bare TypeError: terminated.

    The default is unchanged at 120 s, and the precedence table for signal / requestTimeoutMs / timeoutMs is untouched. One consequence worth planning for: a legitimately large body (a multi-MB /v1/fetch document) now counts against that same budget, where previously only the wait for headers did. (#2546)

  • The README's webhook-retry sample used an invalid delivery id prefix (whd_) that its own branded type rejects. The sample is corrected, and a new test compiles every TypeScript sample in the README against the SDK's sources so a non-compiling sample can no longer ship. (#2547)

Fixed (BREAKING at compile time)

  • fetch.get no longer accepts the eleven markdown-tuning options (markdown_mode, markdown_query, markdown_links, markdown_images, with_links_summary, with_images_summary, markdown_compact, markdown_filter_images, markdown_include_media, markdown_include_warnings, markdown_include_stats). The spec's GET /v1/fetch declares only url, format, country, timeout_ms and cache_ttl, so those options were silently dropped by the query builder — a caller asking for markdown_mode: 'llm' compiled and received default article-mode markdown. They are now compile errors instead of silent no-ops; use fetch.create (POST), which accepts and serializes all of them unchanged. FetchGetOpts is now derived from the generated getFetch query type, so future spec drift fails tsc --noEmit rather than reaching production. (#2539)

[0.3.0] - 2026-07-21

Removed (BREAKING)

  • The deprecated RapidCrawl, RapidCrawlError, and RapidCrawlConnectionError exports and the RapidCrawlOptions type alias (kept as a one-minor-version bridge in #2421) are removed. Import the ByteKit* names instead. (#2473)
  • The RecordingId type export is removed. Recordings are deactivated (/v1/recordings returns 410), so the branded id had no reachable use. Breaking in theory only — unused in practice. (#2481)

Fixed

  • fetch.create / fetch.get no longer throw a bare SyntaxError on every non-JSON body. /v1/fetch returns the fetched content AS the response body (metadata in X-Fetch-* headers), but both methods routed through the JSON-parsing transport. They now use a raw-capable transport and resolve to a typed FetchResult — raw body, byte-exact bytes, plus id, creditsCharged, cache, custom, contentType, and headers parsed from the X-Fetch-* metadata. (#2466)
  • A trailing slash on the constructor baseUrl (e.g. "https://api.bytekit.com/") no longer produces double-slash request URLs — it is normalized once, preserving a path-carrying baseUrl. (#2481)
  • SearchResult.date is now optional (string | null | undefined): images search results omit the field entirely. (#2481)

Added

  • New usage resource (get, daily, byEndpoint — GET /v1/usage, /v1/usage/daily, /v1/usage/by-endpoint) and webhooks resource (list, retry — GET /v1/webhook-deliveries, POST /v1/webhook-deliveries/{id}/retry), with their request/response types exported. (#2471)
  • ScrapeOpts and FetchOpts now cover the full documented request field set (10 previously missing ScrapeRequest fields; 11 markdown-tuning FetchRequest fields), plus the 'text' markdown-links mode. A compile-time drift guard keeps them aligned with the generated OpenAPI schema. (#2475)
  • monitors, sitemap, bulk, and fetch.bulk methods are now fully typed from the generated OpenAPI operations (previously Record<string, unknown> in / Promise<unknown> out); their request/response types are re-exported. (#2481)
  • monitors.list accepts typed query params and returns a typed response (ListMonitorsParams / ListMonitorsResponse, both exported). (#2480)

Changed (BREAKING)

  • fetch.get now rejects a custom option at compile time: GET /v1/fetch has no custom query param, so the value was silently dropped before. Use fetch.create (POST) to echo a custom payload. (#2481)
  • FetchOpts.format is narrowed from string to the spec union 'markdown' | 'html'; off-spec values now fail to compile. (#2475)

[0.2.2] - 2026-07-16

Changed

  • Re-release of 0.2.1. The 0.2.1 artifacts never reached npm — the prod publish's publish-sdk step failed on a CI test-isolation flake (cjs-require.test.ts rebuilding the shared dist), which is now fixed. No consumer-facing changes beyond 0.2.1.

[0.2.1] - 2026-07-16 [NEVER PUBLISHED]

Never published to npm. This version was cut in-tree but no @hunt-labs/[email protected] tarball exists — npm i @hunt-labs/[email protected] fails. Everything below shipped to users in 0.2.2, which is the first published release containing it. Reconciled by issue #2680; the release-drift CI gate now blocks a new phantom heading from being introduced.

Changed

  • Primary exports renamed RapidCrawl→ByteKit, RapidCrawlError→ByteKitError, RapidCrawlOptions→ByteKitOptions. The old names remain as deprecated, identity-preserving aliases (RapidCrawl === ByteKit), so existing imports keep working. (#2421)

Fixed

  • request() now returns undefined on 204 No Content / empty-body responses instead of throwing a JSON-parse error — account.apiKeys.revoke() now resolves as its Promise<void> contract declares. (#2450)

Removed

  • client.recordings (the /v1/recordings endpoint is permanently 410 deactivated). (#2421)

[0.2.0] - 2026-07-15

This is a breaking release. Four runtime behavior changes affect code written against 0.1.0 — read each one before upgrading.

Changed (breaking)

  • Default client-side timeout (120s). Requests now abort after 120000 ms by default (DEFAULT_TIMEOUT_MS in src/client.ts); 0.1.0 had no client-side timeout and would hang indefinitely. A slow request now rejects with RapidCrawlConnectionError carrying code: 'timeout'. Override per client via new RapidCrawl({ apiKey, timeoutMs }).
  • Network failures throw RapidCrawlConnectionError. Every fetch() rejection is now wrapped by toConnectionError() and re-thrown as RapidCrawlConnectionError. 0.1.0 surfaced undici's bare TypeError: fetch failed. Code that matched on a bare TypeError to detect a network failure will no longer match — catch RapidCrawlConnectionError instead.
  • Constructor throws on an empty apiKey. new RapidCrawl({ apiKey: '' }) (or a missing key) now throws a TypeError immediately instead of silently sending Authorization: Bearer undefined. Ensure apiKey is populated (e.g. from process.env.BYTEKIT_API_KEY) before constructing the client.
  • Request-body types narrowed to the OpenAPI schema. Request payloads are now typed from spec-derived types generated from the OpenAPI schema (src/generated/api.d.ts) rather than Record<string, unknown>. Payloads that previously compiled by virtue of the loose index type may now be rejected by the type checker; align them with the documented request shape.

Changed

  • Package renamed to @hunt-labs/bytekit-sdk (the @bytekit npm scope was unobtainable). The prior rapidcrawl/@rapidcrawl/sdk name remains installable at its last version but is deprecated. Corrected in 0.7.1 (see above): that claim does not hold — both names 404 on npm and there is no published version left to deprecate.
  • Generated API types (dist/generated/api.d.ts) are now shipped in the tarball so consumers get the typed request/response surface without regenerating from the spec.

Added

  • Publish-ready packaging metadata: types-first exports, engines, keywords, homepage, bugs, README.md, LICENSE, and an MIT license field.

[0.1.0] - 2026-04-24

Added

  • Initial release of the ByteKit TypeScript SDK — a typed client for the ByteKit API (scrape, screenshot, search, and related endpoints), published under the original @rapidcrawl/sdk name.

bytekit-sdk

Python SDK, published on PyPI. Import name: bytekit.

Unreleased

0.9.2

Release cut 2026-10-02.

Changed

  • Account terms_version docstring describes the new onboarding contract (#5055) — an empty string now means the account has not accepted the Terms yet (bootstrap no longer records acceptance; POST /v1/account/accept-terms does). Docstring only, no model field changes. Regenerated from docs/api/openapi.yaml.

0.9.1

Release cut 2026-09-30.

Changed

  • Bulk job models add total_billed_bytes (#5021) — new Union[None, int] on the bulk job response models (CreateBulkResponse202, GetBulkResponse200/202, DeleteBulkResponse200, and the create/get models of /v1/scrape/bulk and /v1/fetch/bulk) and on BulkCompletedWebhook: the bandwidth billed for the job, in bytes, after multipliers. None on the create response and for a bulk that predates the field. total_credits_charged and the bulk item credits_charged are deprecated in the spec and report 0 under the dual and bytes billing modes. Additive, patch-level. Regenerated from docs/api/openapi.yaml.

0.9.0

Release cut 2026-09-24.

Changed

  • CreateSearchResponse502 is renamed CreateSearchResponse503 (#4891, BREAKING), and CreateSearchResponse502Error is renamed CreateSearchResponse503Error. The gateway now answers a search-provider failure with HTTP 503 instead of 502, because Cloudflare replaced the 502 body with its own plain-text page and the search_provider_error code never reached the caller. The body shape is unchanged. Regenerated from docs/api/openapi.yaml.
  • get_fetch and post_fetch no longer declare a 502 response (#4891). The gateway relays a target's 502 or 504 at HTTP 503, with the true upstream status kept in the body's error.http_status, and the 503 branch already parsed that Error envelope. A 502 is now undocumented, so the default client raises bytekit.errors.UnexpectedStatus for it. Regenerated from docs/api/openapi.yaml.

0.8.2

Release cut 2026-09-23.

Changed

  • AccountSessionResponse.onboarding_completed (#4780) — new bool: whether the account has left dashboard onboarding. Additive, patch-level. Regenerated from docs/api/openapi.yaml.

0.8.1

Release cut 2026-09-18.

Fixed

  • ListBulkScreenshotsResponse200ItemsItemType0.billing (#4582) — now typed Union[BillingNode, None, Unset]. The gateway now returns billing: null on every bulk item whose status is not completed, so an item keeps the same keys on every poll; the previous model passed that null to BillingNode.from_dict and failed to parse the page. A completed item still carries its node, and a completed item from before the node existed still omits the key (UNSET). Regenerated from docs/api/openapi.yaml.

0.8.0

Release cut 2026-09-14.

Added

  • ScrapeRequest.solve_challenge (#4448) — optional bool, default False on the server. When True, the scrape attempts to clear a Cloudflare challenge; success is not guaranteed. It is billed as a normal browser render. Regenerated from docs/api/openapi.yaml; a request built without it serializes with no solve_challenge key, so existing calls are unchanged on the wire.
  • GetUsageResponse200.credits_scrape and GetUsageDailyResponse200DataItem.credits_scrape (#4452) — required int: credits drawn under the scrape usage source, such as /v1/schema extractions. The gateway always charged this bucket but never returned it, so the published per-source credit fields could not add up to credits_used. Regenerated from docs/api/openapi.yaml. The field is required, so a response from a gateway that predates it raises bytekit.errors.UnexpectedStatus (with the KeyError as its cause) on the default client, and returns None when the client is built with raise_on_unexpected_status=False.

Changed

  • ScrapeErrorEnvelope and create_scrape docstrings (#4454) — no longer say an inline POST /v1/scrape failure carries id: null. The gateway now records a sync scrape that ends upstream_error or blocked as a failed scrape job, so the failed envelope carries that job's sc_ id and GET /v1/scrape/{id} returns the same envelope. Regenerated from docs/api/openapi.yaml; model fields and operation signatures are unchanged.

0.7.0

Release cut 2026-09-12.

Added

  • BulkCompletedWebhook / BulkCompletedWebhookItemsItem (#4340) — the body ByteKit POSTs to your endpoint with X-ByteKit-Event: bulk.completed, importable from bytekit.models. The schema has always been in docs/api/openapi.yaml, but openapi-python-client emits models only for schemas reachable from a generated OPERATION body, and no operation $refs this one — it describes a request we make to your server, not a response you read from ours. The generator therefore emitted nothing and every Python caller re-derived the body type by hand. The model is hand-written under scripts/hand_written/ and installed after each regeneration; a spec-reading drift test (tests/test_bulk_completed_webhook_model.py) fails the build if the spec's required keys and the model's diverge.

Changed

  • Bulk list-item models are now a two-member union, and the old single item class is gone (#4340, BREAKING). A terminal bulk item arrives on the wire in one of two disjoint shapes: an ordinary rehydrated item carrying status, or the bounded 16-MiB fallback record carrying terminal_outcome. The spec described both with one flat object, so the generated model made every terminal signal optional and a caller could not tell which field to read. GET /v1/fetch/bulk/{id} (200 and 202) and GET /v1/bulk/{id}/screenshots now declare it as oneOf: [ordinary, fallback] — matching how GET /v1/scrape/bulk/{id} already declared it — and the generator emits one class per alternative:

    RemovedReplaced by
    GetFetchBulkResponse200ItemsItemGetFetchBulkResponse200ItemsItemType0 (ordinary) and GetFetchBulkResponse200ItemsItemType1 (fallback)
    GetFetchBulkResponse202ItemsItemGetFetchBulkResponse202ItemsItemType0 / …Type1
    ListBulkScreenshotsResponse200ItemsItemListBulkScreenshotsResponse200ItemsItemType0 / …Type1

    The per-field enum classes follow the same rename (…ItemsItemStatus → …ItemsItemType0Status, …ItemsItemTerminalOutcome → …ItemsItemType1TerminalOutcome).

    Upgrading: items is now list[…Type0 | …Type1]. Code that read item.status still works on the ordinary alternative; branch on the member you got — isinstance(item, …ItemsItemType0) — or read the terminal signal as getattr(item, "status", None) or item.terminal_outcome. terminal_outcome outranks status when a record carries it, and neither is ever absent.

0.6.0

Release cut 2026-09-11.

Changed

  • create_scrape now raises on a terminal scrape failure, even though it arrives on a 200 (#1870). The synchronous POST /v1/scrape used to answer an unreachable host, a refused connection, or a bot-protection block that defeated every proxy tier with a 503. It now answers 200 carrying the canonical failed envelope, because a failed fetch is a result rather than a transport error. Left alone, the generated union would have parsed that envelope and RETURNED it on the same arm a caller reads .formats from, so the most common first-call shape — "it returned, so it worked" — would have silently succeeded on a failed scrape. create_scrape.sync / .sync_detailed / .asyncio / .asyncio_detailed therefore raise errors.UnexpectedStatus with .status_code 200 and .code / .message read off the envelope's error; under raise_on_unexpected_status=False they return None, as for any other unusable body.

    The carve-out is that one call site. get_scrape still RETURNS the identical envelope as a typed ScrapeErrorEnvelope — polling a job that failed is a successful poll — and every documented 4xx/5xx on create_scrape still returns its typed Error model.

    Upgrading: if you inspected the result of create_scrape for status == "failed", move that handling into an except errors.UnexpectedStatus block. Code that already caught UnexpectedStatus needs no change.

0.5.4 [NEVER PUBLISHED]

Release cut 2026-09-06.

Never published to PyPI. Publication is triggered by CI on main from the version pyproject.toml declares, and 0.6.0 (#1870) superseded 0.5.4 in the tree before any release reached main — no bytekit-sdk 0.5.4 artifact exists, so pip install bytekit-sdk==0.5.4 fails. Everything below shipped to users in 0.6.0, the first published release containing it.

Added

  • Bulk fetch, scrape, and mixed-result models now publish cursor pagination, stable item identities, whole-job counters, exact mixed artifact keys, artifact warnings, bounded webhook recovery metadata, and the artifact_unavailable error contract (#3717). Generated bulk read functions accept optional cursor, limit, and status parameters.

Fixed

  • Corrected the artifact-URL fallback docstrings on list_bulk_screenshots (image_url, content_url, markdown_url) and the /v1/logs row model (artifact_url) (#3717). Those fields were documented as null-plus-artifact_unavailable when no signature can be minted; the servers have always returned the row's stored value unchanged instead — None while nothing is stored, and otherwise the private unsigned storage URL, which answers 401 — and raise no warning. get_fetch_bulk does null the field and report artifact_unavailable, and its docstring is unchanged. No server behaviour changed and no model moved; only the documentation on them did.

Changed

  • Regenerated GetUsageResponse200 against the corrected topup_balance_bytes description (#4251). The field's docstring read "Remaining purchased top-up bytes", which contradicted its own "Equal to topup_bytes" clause: the value covers purchased top-up lots and the one-time starter grant, which is what the endpoint has always returned. The clause recording that the figure is already net of consumed overage, and the "Present only in dual/bytes billing mode" clause, are both unchanged and still true. Types are unchanged and no server value changed; only the documentation on them moved.
  • Regenerated ScrapeSuccessEnvelope against the clarified cache_age_s description (#4237). The field's docstring now records that GET /v1/scrape/{id} reports cache: hit without cache_age_s: the age is a property of the replay that happened at completion time, is not persisted on the row, and cannot be derived honestly afterwards. The one-directional "Present only on cache: hit" clause is unchanged and still true. Types are unchanged; only the documentation on them moved.
  • Regenerated ScrapeSuccessEnvelope against the completed cache presence list (#4203). The field's description now names GET /v1/scrape/{id} as a third delivery path, alongside the synchronous response and webhook delivery: a polled scrape now reports the cache disposition its row recorded, and omits the field — rather than defaulting it to miss — for a row that recorded none. Types are unchanged; only the documentation on them moved.
  • Regenerated MonitorCreateRequestScrapeOptions against the corrected scrape_options documentation (#4024). The description no longer lists headers among the accepted fields — POST /v1/monitors never honoured a caller-supplied headers option and now rejects it with the same 400 every other capture endpoint returns for the field. Types are unchanged; only the documentation on them moved.

0.5.3

Release cut 2026-09-05.

Changed

  • Regenerated ScrapeQueuedEnvelope against the corrected status_url documentation (#3925). status_url is a root-relative path, never an absolute URI: join it onto the API base URL you called (https://api.bytekit.com, no trailing slash) to get the URL to poll. The spec had declared it format: uri with an absolute example, which no server has ever emitted; the emitted value is unchanged, so no client behaviour moves.

0.5.2

Release cut 2026-09-03.

Changed

  • Breaking (API): the scrape formats vocabulary is now raw | markdown | links | images (#3775). ScrapeRequestFormatsItem drops the two retired members and gains RAW, and the response model exposes ScrapeFormats.raw in place of the old raw-HTML and cleaned-HTML fields. The server accepts no alias, so a request naming a retired value returns 422:

    body = ScrapeRequest(url="https://example.com", formats=[ScrapeRequestFormatsItem.RAW])
    resp = create_scrape.sync_detailed(client=client, body=body)
    print(resp.parsed.formats.raw)

Fixed

  • 402 Payment Required now returns a typed Error instead of raising UnexpectedStatus on create_scrape, get_fetch, post_fetch, create_fetch_bulk, create_scrape_bulk, create_bulk and create_sitemap (#3784). Those operations enforce quota but never declared 402 in the spec, so the generated client had no arm for it while create_search and create_screenshot did — the same exhausted account raised on one call and returned a parsed error on the next. Regenerated from the corrected spec:

    resp = create_scrape.sync_detailed(client=client, body=body)
    if resp.status_code == 402:
        print(resp.parsed.error.code)  # quota_exceeded | spending_cap_reached | …

Changed

  • The generated events field docstring on ScrapeRequest now documents the widened server default (queued, completed, failed) and that a caller may send an explicit subset to receive fewer deliveries (#3777). Types/docs only — this regeneration followed a server-side default change; the SDK carried no logic tied to the old default.

0.5.1

Release cut 2026-08-25.

Fixed

  • GetScrapeBulkResponse200 now exposes the items a completed scrape-bulk poll returns (#3632). get_scrape_bulk parsed the terminal 200 into a model with no items field, so a create-then-poll flow surfaced typed envelopes on every non-terminal poll and lost them on the terminal one, where the results actually live:

    polled = get_scrape_bulk.sync(id=job.id, client=client)
    for item in polled.items:  # typed ScrapeSuccessEnvelope / ScrapeErrorEnvelope / ScrapeQueuedEnvelope
        ...

    Server behavior is unchanged — the field was always on the wire.

  • Polling a bulk job that is still running no longer raises UnexpectedStatus: 202 (#3533). get_scrape_bulk, get_fetch_bulk and get_bulk return HTTP 202 with the job envelope while the job is non-terminal and 200 once it is terminal, but only the 200 was documented, so the first poll of the obvious create-then-poll workflow raised (or returned None with raise_on_unexpected_status=False). Each of the three now parses the 202 into its own typed model — GetScrapeBulkResponse202, GetFetchBulkResponse202, GetBulkResponse202 — which carries the same fields as the matching 200 model:

    job = create_scrape_bulk.sync(client=client, body=body)
    polled = get_scrape_bulk.sync_detailed(id=job.id, client=client)
    # polled.status_code is 202 while processing, 200 once terminal; .parsed is typed in both

    Server behavior is unchanged — this is a client-side parsing fix.

Changed

  • GetUsageResponse200.topup_balance_bytes now documents itself as the remaining purchased top-up bytes, already net of any overage consumed against them (#3459). Documentation only: the field's name, type and wire format are unchanged. The server previously subtracted the consumed overage a second time when computing this value, so the number a client reads is larger than before for accounts that ran past their included allowance.

0.5.0

Release cut 2026-08-20.

Removed

  • BREAKING: the cookies and headers request-body fields are gone from every capture endpoint (#3242) — /v1/scrape, /v1/scrape/bulk, /v1/screenshots, /v1/recordings and /v1/schema. A request that carries either field is rejected with 400 validation_error; it is not ignored, and there is no deprecation window, feature flag, or compatibility header. ScrapeRequest.cookies, ScrapeRequest.headers, the matching ScreenshotRequest fields and the generated Cookie model are removed from the package.

    # before
    client.scrape(url=url, cookies=[Cookie(name="session", value="abc")])
    # after — the keyword no longer exists; the same body over raw HTTP is a 400 validation_error
    client.scrape(url=url)

    ByteKit continues to manage its own cookie jar and request headers internally; what was removed is the caller's ability to inject or override them.

Changed

  • fetch now parses a 503 response instead of treating it as unexpected (#3372). /v1/fetch answers 503 on two conditions and both carry the ordinary Error envelope, so the generated client returns an Error for them rather than raising UnexpectedStatus (or returning None) the way an undeclared status is handled. Read error.code to tell the two apart, because they disagree about error.http_status:

    • upstream_error — the target returned HTTP 502, relayed at 503 because Cloudflare replaces a clean origin 502 with its own error page. error.http_status stays 502, the true upstream status, and is the only surviving record of what the target answered.
    • internal_error (Fetch service unavailable.) — the fetch service was unreachable. There is no upstream, so error.http_status reports 503, matching what was served.

    Nothing else about the call changes, and every other documented /v1/fetch status is unaffected.

0.4.0

Release cut 2026-08-19.

Removed

  • BREAKING: billing_multiplier and billed_bytes are gone from ScrapeSuccessEnvelope (#3143). Read billing instead — a BillingNode that names every factor applied to the charge, the multiplier they compose to, and the resulting billed bytes:

    # before
    spend = res.billed_bytes            # flat, base-endpoint only
    # after
    spend = res.billing.billed_bytes    # guard on isinstance(res.billing, BillingNode)
    why = res.billing.factors           # [BillingFactor(name="endpoint", value=1.5, reason="scrape_md"), ...]

    This is a correction, not a rename: the two surfaces reported different numbers for the same charge. The flat pair carried the BASE-endpoint composition, so a markdown cache hit read 0.5 / 400 while the real charge — and the X-Billing-Multiplier header — was 0.75 / 600; an html-only miss read the per-format credits estimate 1.5 / 150 against an endpoint factor of 1.0 / 100. The flat pair also never carried the clean_markdown x3 surcharge, so on a successful clean it under-reported the amount billed by two thirds. billing.multiplier composes every applied factor behind exactly one rounding, so it is the number that predicts your invoice.

    billing is UNSET on resources that completed before the billing node existed. A pre-node row omits the field rather than reporting a zero node, because a zero node would assert that a real past charge was zero. Test with isinstance(res.billing, BillingNode), never with a truthiness check — a legitimate zero-byte charge is a value, not an absence.

    Two endpoints deliberately carry no node: /v1/search is billed per credit and reports credits_used, and a single /v1/fetch reports its charge through response headers only. The /v1/logs row field named billing_multiplier is a different field and is not the envelope field this release removed. It carries the BASE settled credits_charged factor, not the composed real-endpoint value the X-Billing-Multiplier header reports — a markdown cache hit reads 0.5 on the logs row and 0.75 on the header.

Changed

  • Generated ScrapeErrorEnvelope documentation now describes terminal asynchronous scrape.failed webhook delivery (#1856). No runtime SDK behavior or API shape changed.

Fixed

  • GET /v1/logs entries now carry billing_multiplier, as a JSON number (#3022). The gateway had been serializing that value under an undeclared key credits, and as the raw Postgres NUMERIC(10,2) string ("1.50"), while the OpenAPI spec declared billing_multiplier as a number.

    Breaking at the wire level: the undeclared credits key is gone from every /v1/logs entry. Code reading credits off a logs response must read billing_multiplier. No SDK method exposes this endpoint, so nothing inside this package changed.

  • change_threshold on MonitorResponse/MonitorCreateResponse and change_pct on CaptureResponse are now typed float, not str (#3021). Both are Postgres NUMERIC(5,2) columns; the gateway passed the raw driver string through unconverted ("5.00" / "7.25") while the OpenAPI spec's request schemas already declared number — only the response schemas had been (incorrectly) edited to match the buggy string output. This release fixes the gateway serializer and corrects the two response schemas, so spec and runtime now agree.

    Breaking at the wire level: bytekit.api.monitors.create_monitor, bytekit.api.monitors.get_monitor, bytekit.api.monitors.list_monitors, and bytekit.api.monitors.update_monitor now return float for change_threshold instead of str (screenshot-type monitors only, create_monitor/update_monitor included since both also serialize through the same MonitorCreateResponse/MonitorResponse schemas); bytekit.api.monitors.list_monitor_captures now returns float for change_pct instead of str. change_pct is None (not 0.0) when there is no prior capture to diff against, unchanged from before. Only the generated model description changed inside this package.

0.3.9

Release cut 2026-08-11.

Changed

  • The declared httpx floor is raised from >=0.24.0 to >=0.28.0 (#2998). The SDK's own tests assert the exact request and response bytes it puts on the wire, and those bytes changed in httpx 0.28.0, which switched to compact JSON separators (b'{"a":"b"}'; every release up to and including 0.27.2 emits b'{"a": "b"}'). The old floor therefore promised an install the suite cannot run green: on httpx 0.24–0.27 the dependency resolves, and 12 assertions then fail with a message naming neither httpx nor a version. Raising the floor was chosen over loosening the assertions, because the compact form is the wire format the SDK actually emits and pinning it is the reason those tests exist. No runtime behaviour of this package changes — only what pip will resolve. httpx 0.28.0 declares requires-python >=3.8, so this excludes no Python version the SDK supports (requires-python = ">=3.10").

  • Generated vocabularies follow the spec's corrected error and format enums (#2959). ScrapeErrorEnvelopeErrorCode no longer carries QUOTA_EXHAUSTED: the API has emitted quota_exceeded since #1539 and no gateway throw site produces quota_exhausted, so this removes a member that has never arrived over the wire — not one that stops arriving. Code matching on it was matching a branch that could not be reached; the live code for an exceeded quota is quota_exceeded. INVALID_URL, which the API does emit, is unchanged.

    CreateBulkBodyDefaultsFormatsItem and CreateBulkBodyItemsItemFormatsItem rename RAWHTML = "rawHtml" to RAW_HTML = "raw_html", matching ScrapeRequest and /v1/scrape/bulk. /v1/bulk accepted only the internal spelling and the spec documented that; it now accepts both, so the one endpoint that disagreed with the rest of the surface no longer does. rawHtml keeps working on the wire.

Fixed

  • An invalid type or date_range passed to client.search(...) now raises an error you can act on (#2958). search(type="bogus") raised a bare ValueError: 'bogus' is not a valid CreateSearchBodyType straight out of a generated enum constructor — it named a class you never imported, it did not say which values would have worked, and the method's Raises: section listed only UnexpectedStatus, so the error was undocumented as well as unhelpful. The message now names the parameter, the value it rejected and every value it accepts, and it is raised before any request is sent, so a typo can never be billed.

    The type hints for both arguments narrowed from Optional[str] to a Literal union with the generated enum class, which makes a wrong CONSTANT a type error at the call site rather than a runtime surprise. This is hints-only: any str is still accepted at runtime, so code that computes the value — from argv, a config file, a database column — keeps working unchanged. Enum members (CreateSearchBodyType.IMAGES) type-check as before.

  • A successful 200 carrying an extraction-quality warning no longer fails to parse (#2948). ScrapeWarningCode was generated from a spec enum that listed 10 of the 12 codes the server emits; low_quality_extraction and fragment_menu_stripped (added server-side by #1701) were missing. ScrapeWarning.from_dict builds the code with ScrapeWarningCode(...), which raises ValueError on an unknown member, and #2684's bytekit_parse guard then turned that into the two-tier contract — so a successful, billed 200 reached the caller as errors.UnexpectedStatus under the default raise_on_unexpected_status=True, and as response.parsed is None under the opt-out. Both codes are now in the spec and the enum, so such a response parses normally. Callers who added a defensive except UnexpectedStatus around scrape calls for this reason can drop it. The other 10 codes are unaffected.

Changed

  • Connection failures are documented. Every Raises: section named httpx.TimeoutException and stopped there, and the README's error-handling section did the same. Its transport sibling httpx.ConnectError — DNS failure, connection refused, TLS handshake failure — is just as reachable and is caught nowhere, so except (UnexpectedStatus, httpx.TimeoutException) written from the documented surface left an uncaught exception waiting for the first flaky network. Both client-class docstrings, search()'s Raises: section and a new README Transport errors table now name both. No behavior changed: neither exception was ever caught, and neither is now.

  • README: two first-use traps. get_fetch takes url_query=, not url= (it travels as the url query parameter), and AuthenticatedClient.search(...) is synchronous — calling it inside a coroutine blocks the event loop, with create_search.asyncio as the async alternative. Both are now documented with runnable examples.

  • The sdist no longer ships a legacy-brand name. bytekit-sdk==0.3.8 carried one occurrence of the pre-cutover project name — inside a .gitignore this package does not own and its build configuration does not list. hatchling force-includes the nearest ignore file into an sdist, which here means the repository's, so the string reached PyPI while a search of the package itself came back clean. The line is corrected, and the BUILT tarball is now grepped on every test run rather than the source tree. Nothing you install changes; the artifact simply stops advertising a name and a command that no longer exist.

  • A release can no longer be published while its own changelog denies it exists. The 0.3.8 sdist shipped with its ## 0.3.8 section still carrying the cut-time placeholder that denies publication: it was written when the release was cut, was never updated when the release actually happened, and nothing existed to notice. The publish step now refuses to upload a version whose section still carries that placeholder. This entry documents the fix for readers of the artifact; the guard itself lives in the repository. (The placeholder's own wording is deliberately not quoted here — the guard reads this section, and quoting it would trip the very check being described.)

  • Alignment release — the published 0.3.8 artifact already carries this content. 0.3.8 published to PyPI at 2026-08-11T21:18:32.641Z (sdist bytekit_sdk-0.3.8.tar.gz), built from the promotion HEAD, so it contains every change documented under 0.3.8 — including the two review fix-ups to the injected client_surface converter (non-mapping headers= values, and case-insensitive override precedence). Verified by downloading the published sdist and reading src/bytekit/client.py. Nothing about the installed package changes through this alignment — the #2948 enum fix above is a separate, deliberate change riding the same unpublished version.

    This release exists because check-release-drift (#2680) is version-gated and anchors on the commit that introduced a version, not on the commit whose tree was actually published. Those two fix-ups landed after the commit that introduced 0.3.8, so once 0.3.8 reached PyPI the gate read them as post-release drift. Bumping moves the anchor past them. This mirrors the 0.3.6 version-alignment release cut after the 2026-08-10 publish rescue.

0.3.8

Published to PyPI 2026-08-11 (21:18:32.641Z).

Added

  • Every request now carries an X-ByteKit-Client header identifying this SDK and its version — X-ByteKit-Client: sdk_python/<version>, sent by default from both Client and AuthenticatedClient, on the synchronous and the asynchronous httpx client alike. It lets ByteKit tell traffic that came through the Python SDK apart from direct API calls; before this release those were indistinguishable. The version half resolves from the installed distribution's metadata (the same source as bytekit.__version__), so it can never drift from pyproject.toml.

    A caller-supplied header of the same name wins, whether it is passed to the constructor (Client(headers={"X-ByteKit-Client": "…"})) or merged in later (client.with_headers({"X-ByteKit-Client": "…"})) — the marker is a telemetry dimension, not an authentication or authorization control, and ByteKit treats an unrecognized value exactly as it treats a missing one. Unrelated caller headers leave the marker intact.

    The override is matched case-insensitively, as HTTP header names are, so x-bytekit-client and X-BYTEKIT-CLIENT override just as X-ByteKit-Client does — including through httpx.Headers, which lowercases every key it is given. Exactly one X-ByteKit-Client header goes on the wire either way; you will never see your override and the default sent together.

    Nothing else about request construction changed: AuthenticatedClient still stamps its auth header, with_headers still merges caller keys over defaults, the cookies, timeout, verify and redirect settings of both clients are untouched, and headers= still accepts everything it accepted before — a dict, None, a list or tuple of pairs, an httpx.Headers.

Fixed

  • A headers dict you pass to a client is no longer written into by the SDK. Previously the client kept your mapping by reference and stamped Authorization: Bearer <your token> into it when it built its httpx client, so a dict you still held — and might log, reuse for a second client, or share across threads — silently acquired your API key. The client now copies what you pass, so your object is left exactly as you gave it and two clients built from one dict can no longer contaminate each other's credentials. Nothing about the headers the client sends changed.

0.3.7 [NEVER PUBLISHED]

Never published to PyPI. This version was cut in-tree but no bytekit-sdk 0.3.7 artifact exists — pip install bytekit-sdk==0.3.7 fails. Everything below shipped to users in 0.3.8, which is the first published release containing it.

Documentation

  • AccountSessionResponse.has_activity now documents what it actually means. The generated model's attribute docstring previously said the flag was true once the account had "ever recorded billable activity", which read as "ever made a request". It is true only once a request has successfully completed and been charged credits — a request that failed does not flip it, even when its bandwidth was billed (issue #2890). The docstring also now states the all-time (never windowed) scope and the up-to-one-rollup lag, so a caller polling it after a first request knows False can mean "not visible yet" rather than "never made a request". Regenerated from docs/api/openapi.yaml; the attribute's TYPE is unchanged (has_activity: bool), so this release is source-compatible in both directions.

0.3.6

Documentation

  • No README change ships in 0.3.6. The README's Error handling documentation reached PyPI inside the published 0.3.5 artifact (2026-08-10 publish rescue) and stays documented under the 0.3.5 heading, which is the version whose wheel actually carries it.

Changed

  • Version-alignment release after the 2026-08-10 publish rescue: the published 0.3.5 artifact was built from the main promotion HEAD, which already contained every change documented under the 0.3.5 heading — but that commit postdates the one that introduced the 0.3.5 version string, so the check-release-drift gate (#2680) correctly reported shipped drift on every subsequent CI run (#2829). 0.3.6 realigns manifest ↔ registry ↔ version-introducing commit. No consumer-visible change relative to the published 0.3.5.

0.3.5

Breaking

  • AuthenticatedClient takes no positional arguments. token, prefix and auth_header_name are now keyword-only, joining base_url, timeout, raise_on_unexpected_status and the rest — so AuthenticatedClient(...) accepts keywords only. Before 0.3.0 the first positional argument was base_url, which made AuthenticatedClient("https://api-stg.bytekit.com", "sk_live_…") a documented call. When 0.3.0 made base_url a defaulted keyword argument, that call stopped meaning what it said and started binding the URL to token and the API key to prefix — sending Authorization: sk_live_… https://api-stg.bytekit.com to the default host, https://api.bytekit.com, with nothing raised and nothing warned. It now raises TypeError. Migration is mechanical: AuthenticatedClient(base_url="https://api-stg.bytekit.com", token="sk_live_…"). Keyword construction — the form the README, the docs site and every example already use — is entirely unchanged, including prefix=/auth_header_name= overrides and the with_headers/with_cookies/with_timeout helpers. The unauthenticated Client was already keyword-only and is unaffected.

Fixed

  • A schema-drifted response body no longer crashes with a bare KeyError. A generated model's from_dict trusts the required keys the OpenAPI spec declared when the client was generated, so a server that renamed, dropped or retyped a field leaked the generator's own internal exception straight to the caller — KeyError: 'schema_version' from create_scrape, KeyError: 'period_start' from get_usage, KeyError: 'data' from list_monitors, plus TypeError/ValueError for non-object and nested-drift bodies. It happened on a documented status (a 200/202), in both raise modes, and no row of the documented error table covered it, so a caller who handled that table faithfully still had no branch for it. Schema drift now takes exactly the same two arms as a non-JSON body: errors.UnexpectedStatus (carrying .status_code and the raw .content, with the original parse failure attached as __cause__) in the default mode, and None under raise_on_unexpected_status=False. Deliberately no new exception type — a new class would be one every existing except errors.UnexpectedStatus silently fails to catch, i.e. the same untyped crash in a different costume. Bodies that match the documented schema parse exactly as before, and a genuinely optional key that is absent still yields UNSET.
  • client.search() never returns a silent None. Its 200 body is parsed through the same shared guard the generated operations use, which honors raise_on_unexpected_status — so an opt-out client got None handed back from a non-JSON 200 or a schema-drifted one, contradicting both the method's non-Optional return annotation and its own docstring. search() now raises errors.UnexpectedStatus in both raise modes. raise_on_unexpected_status configures the generated two-tier contract, whose signatures are Optional[...]; search() was never part of it — it already raises on documented error statuses that the generated create_search returns as a typed Error model. A well-formed 200 still returns the parsed CreateSearchResponse200 in either mode.
  • A non-IANA HTTP status no longer crashes every operation with a bare ValueError. Each generated operation built its Response with status_code=HTTPStatus(response.status_code) evaluated before the response was parsed, so any status outside Python's http.HTTPStatus — Cloudflare's 520–530 family, nginx's 499 — raised ValueError: 520 is not a valid HTTPStatus from inside the SDK. This happened in both raise modes: the default raise_on_unexpected_status=True never got to raise its errors.UnexpectedStatus, and raise_on_unexpected_status=False, which promises None, raised the ValueError too. api.bytekit.com is served through Cloudflare, so this was exactly the edge/CDN failure class the typed-error contract exists for. All 29 operations now behave as documented: errors.UnexpectedStatus (carrying .status_code and the raw .content) in the default mode, a Response with .parsed is None in the opt-out mode. The hand-written client.search() wrapper was never affected and is unchanged.

Changed

  • Response.status_code is a plain int for a status http.HTTPStatus does not know. This is the contract decision behind the fix above, recorded here because it is the one observable difference in the typed surface. Concretely:

    • Every status the stdlib enum does know — documented by the API or not, 200 as much as 503 — still comes back as a real http.HTTPStatus member, so response.status_code.phrase, .name, and identity comparisons such as response.status_code is HTTPStatus.OK keep working exactly as before.
    • A status the enum does not know (499, 520, 521, 522, 530, …) comes back as the plain int the wire carried. int comparisons (response.status_code == 520, >= 500) work; enum-only attributes (.phrase, .name) do not exist on it. The alternatives were rejected deliberately: a synthetic IntEnum member would invent a .phrase/.name no registry backs, and raising before the Response is constructed is the bug itself — it is precisely what raise_on_unexpected_status=False promises not to do.
    • The declared annotation on Response.status_code is unchanged (HTTPStatus), so this release adds no type errors to existing consumer code; widening it to Union[HTTPStatus, int] would make .phrase/.name a type error on every response, including the IANA ones, which is a far larger break than the residual it would close. Reach for .phrase/.name only after an isinstance(..., HTTPStatus) check if you handle edge statuses off a non-raising client.

    In the default (raising) mode this is largely invisible: a non-IANA status raises errors.UnexpectedStatus before any Response is returned, and UnexpectedStatus.status_code has always been a plain int.

Documentation

  • The "Error handling" example in the README now requests formats=[…MARKDOWN] explicitly. Run literally, its success arm printed <bytekit.types.Unset object at 0x…>: the snippet built ScrapeRequest(url=...) with no formats, so the server applied its raw_html default and result.formats.markdown was legitimately unset. The same class of defect 0.3.3 fixed in the quick start, in the block directly beneath it. The README's error table gains the schema-drift row, a note that search() is outside the two-tier contract, and a migration note for the keyword-only constructor.

0.3.4

Breaking

  • CreateBulkBody, CreateBulkBodyDefaults, CreateBulkBodyItemsItem, and CreateFetchBulkBodyUrlsItemType1 lost their additional-properties mapping API. docs/api/openapi.yaml's /v1/bulk request schema declares additionalProperties: false on the top-level body, items[], and defaults (issue #2641), and /v1/fetch/bulk's per-item override schema inside urls[] declares the same (issue #2648), matching the gateway's own .strict() validators. The code generator responds to additionalProperties: false by omitting the catch-all mapping it otherwise attaches to every generated model, so these four models no longer expose additional_properties, the additional_keys property, __getitem__, __setitem__, __delitem__, or __contains__. Code that did body["some_key"] = value, "some_key" in body, or read body.additional_keys on any of these four models will now raise AttributeError/TypeError instead. Use the model's declared attrs fields directly instead (e.g. CreateBulkBodyItemsItem(url=..., type=...), CreateFetchBulkBodyUrlsItemType1(url=..., format_=...)); the server itself now rejects unrecognized keys on these endpoints with 422 validation_error, so the removed escape hatch could never have reached the API successfully anyway. This is scoped to the four create_bulk_body* / create_fetch_bulk_body_urls_item_type_1 models — no other generated model is affected.

  • If you are upgrading from 0.3.2 or earlier, note that 0.3.3 also carried a breaking model removal that went undocumented at the time — see the "Breaking" entry under 0.3.3 below (CreateFetchBulkBodyMetadata and CreateScrapeBulkBodyItemsItemMetadata).

0.3.3 [NEVER PUBLISHED]

Never published to PyPI. This version was cut in-tree but no bytekit-sdk 0.3.3 artifact exists — pip install bytekit-sdk==0.3.3 fails. Everything below shipped to users in 0.3.4, which is the first published release containing it. Reconciled by issue #2680; the release-drift CI gate now blocks a new phantom heading from being introduced.

Breaking

  • CreateFetchBulkBodyMetadata and CreateScrapeBulkBodyItemsItemMetadata were removed; both fields now use the shared Metadata model. This landed with the default-stripping traversal fix listed under "Fixed" below (issue #2592) and was not recorded at the time. Once server defaults stopped being materialized into the request-direction schemas, the two per-body metadata objects became structurally identical to the shared Metadata component, and the code generator emits one model per distinct schema — so bytekit.models.create_fetch_bulk_body_metadata and bytekit.models.create_scrape_bulk_body_items_item_metadata no longer exist and importing either raises ModuleNotFoundError. CreateFetchBulkBody.metadata and CreateScrapeBulkBodyItemsItem.metadata are now typed Union[Unset, Metadata]. The wire format is unchanged — the same JSON object is sent either way — so the migration is purely at the import site: from bytekit.models.metadata import Metadata, then Metadata(...) (or Metadata.from_dict({...})) wherever you constructed one of the two removed classes.

Security

  • Path parameters are now percent-encoded. All 13 operations that interpolate an id into their URL (get_scrape, get_screenshot, get_bulk, delete_bulk, list_bulk_screenshots, get_scrape_bulk, get_fetch_bulk, get_monitor, update_monitor, delete_monitor, list_monitor_captures, get_sitemap, retry_webhook_delivery) previously interpolated the value raw, so an id containing /, ?, # or traversal segments could retarget the request to a different path on the same authenticated host — e.g. a DELETE /v1/bulk/{id} becoming a delete against another resource. Values are now encoded with quote(str(value), safe=""). This is not a breaking change for any id the API issues: hex and sc_/ss_/mon_/sm_/bulk_-prefixed ids contain no reserved characters, so the resulting URL is byte-identical to before. No id-format validation was added — the client does not reject id shapes.

Fixed

  • Reusing one client across several asyncio.run(...) calls no longer raises RuntimeError: Event loop is closed. An httpx.AsyncClient's connection pool belongs to the event loop that created it, and the SDK cached its internal async client unconditionally. It now tracks which loop that client was built on and transparently rebuilds it when the running loop changes, re-applying base_url, headers, timeout, httpx_args and authentication. Behavior is unchanged when no loop is running, and a client you supply via set_async_httpx_client(...) is never rebuilt or closed — it stays yours to manage. See the README's new "Clients and event loops" section.
  • The README quick start now prints markdown instead of an Unset placeholder. It requested no formats, so the server applied its raw_html default and result.formats.markdown was legitimately unset. It now asks for formats=[ScrapeRequestFormatsItem.MARKDOWN] — note formats takes enum members, not plain strings.
  • client.search(type="images") no longer crashes with a bare KeyError: 'snippet'. The OpenAPI spec's search result-item schema declared snippet as required while its own description said the key is omitted for images results — a self-contradiction the code generator trusted literally, so CreateSearchResponse200ResultsItem.from_dict unconditionally did d.pop("snippet"). snippet (like date) is now Union[Unset, str], matching the real gateway behavior (web/news results still always carry snippet; images results omit it entirely) — fixed at the spec level (docs/api/openapi.yaml) and regenerated, not patched in the generated model. (Issue #2579.)
  • Server defaults are no longer materialized into request wire bodies/query strings for inline request bodies and $ref'd parameters. scripts/filter-spec.ts's default-stripping traversal previously only followed $ref'd request-body component schemas, so ScrapeRequest/FetchRequest were clean but every other request-direction shape kept baking in server defaults: /v1/search's fully-inline body (type, limit, country, language, date_range), all three bulk endpoints' bodies and their nested Defaults/ItemsItem models, and $ref'd query parameters (GET /v1/fetch's country/timeout_ms/cache_ttl, list_monitors' limit/status, list_webhook_deliveries'/list_monitor_captures' limit, create_screenshot's async). The traversal now walks inline request bodies AND parameters (both operation-level and shared path-item-level, $ref'd or inline) before generation, so an omitted optional field always serializes to nothing and the server's own default applies — caller-supplied values are unaffected and still serialize exactly as given (issue #2592).
  • tests/test_no_materialized_defaults.py rewritten to enumerate every request-direction model and operation parameter programmatically (walking the generated bytekit.api/bytekit.models trees) instead of sampling two known-clean models, so a newly introduced leaker is caught automatically.

Changed

  • The source distribution no longer ships tests/. The bundled suite could not be collected from an unpacked sdist: 12 of its 27 modules require the repository's codegen scripts or the canonical OpenAPI spec, neither of which belongs in a published distribution, and three of them build distributions of the package itself. Nothing importable was removed — the wheel is unchanged, py.typed included, so pip install bytekit-sdk is unaffected.

Internal

  • CI now proves the committed client is reproducible from the canonical OpenAPI spec, and that the generation chain is deterministic (generating twice yields byte-identical output). This closes the gap where a spec change could land with a TypeScript-only regeneration and leave the Python client silently stranded on an older spec.
  • ruff is now pinned to an exact version alongside the code generator. The generator runs ruff over everything it emits, so it — not the generator alone — determines the committed bytes; the generator itself accepts any ruff<0.13, which meant a fresh install could reformat the whole tree. Bumping the pin is a deliberate, reviewed reformat.
  • The PEP 561 py.typed marker is now emitted by the post-generation injector rather than hand-maintained inside the generated tree (the generator omits it under --meta none).

0.3.2

Added

  • bytekit.__version__ reports the installed distribution version, resolved at import time from importlib.metadata (so it can never drift from pyproject.toml's [project].version). When the package is imported from a source tree with no installed distribution it reads 0.0.0.dev0 rather than raising. The dead version override in sdk-python-config.yaml — inert under --meta none and contradicting the real version — has been removed, leaving one authoritative version.

Fixed

  • AuthenticatedClient.search() is now visible to type checkers. It was attached to the class after creation, so despite the package shipping a py.typed marker, mypy and pyright reported "AuthenticatedClient" has no attribute "search". It is now declared as a real method in the class body that delegates to the same implementation: runtime behavior, arguments, return type and raised errors are unchanged — only the static surface is fixed.

0.3.1 [NEVER PUBLISHED]

Never published to PyPI. This version was cut in-tree but no bytekit-sdk 0.3.1 artifact exists — pip install bytekit-sdk==0.3.1 fails. Everything below shipped to users in 0.3.2, which is the first published release containing it. Reconciled by issue #2680; the release-drift CI gate now blocks a new phantom heading from being introduced.

Fixed

  • The "never a bare json.JSONDecodeError" guarantee is now package-wide across all generated operations, not only the AuthenticatedClient.search() wrapper. Every generated operation previously called response.json() unconditionally on a DOCUMENTED status, so a 500/502 serving an HTML load-balancer page raised a bare json.JSONDecodeError from inside the SDK. All operations now parse documented statuses through a shared guard: a non-JSON body raises the typed errors.UnexpectedStatus (or returns None when raise_on_unexpected_status=False), carrying the status code and the raw body.
  • errors.UnexpectedStatus.code / .message are now populated for generated operations too, from the documented Error envelope — previously they were always None outside search(). A string-valued error (the masked search-provider 502) still never enriches, so the upstream provider's identity is never surfaced.
  • Behavior on JSON bodies is unchanged: JSON success bodies still parse into their typed success models, documented JSON error bodies are still returned as typed Error models rather than raised, and raise_on_unexpected_status semantics are untouched. The documented non-JSON success paths — get_fetch/post_fetch's raw-text 200 and delete_monitor's empty 204 — are likewise unaffected.
  • README error documentation corrected. The quick-start example wrapped a call in try/except UnexpectedStatus and then read response.formats.markdown, but a documented 4xx returns a typed Error model without raising — so the documented example failed with AttributeError: 'Error' object has no attribute 'formats'. The README now documents the real two-tier contract, and the snippet is executed by tests/test_readme_examples.py so it cannot silently drift again.

Documentation

  • Removed the contributor-only Publishing, Regenerating, and Development sections from the published README (the PyPI long description) — they described internal CI mechanics and a monorepo-contributor workflow that do not belong on a public package index. The regeneration/development knowledge is preserved in-repo in packages/sdk-python/CLAUDE.md (not shipped in the distribution).
  • Added a ## License section and a LICENSE file, matching the npm sibling packages.

0.3.0

Breaking

  • PyPI distribution renamed bytekit -> bytekit-sdk. PyPI administratively denylists the bare name bytekit, so the package installs as pip install bytekit-sdk. The import name is unchanged — import bytekit still works, and no module path moved. Nothing was ever published under the old distribution name, so there is no migration for existing installs.
  • raise_on_unexpected_status now defaults to True. Undocumented response statuses now raise errors.UnexpectedStatus instead of silently returning None. Callers that relied on the old None-return behavior must pass raise_on_unexpected_status=False explicitly to opt out.
  • base_url is now a keyword-only argument. It gained a default (https://api.bytekit.com), and to satisfy attrs field ordering it became keyword-only. Any caller passing base_url positionally must switch to the base_url= keyword form. (base_url= was already the documented usage.)

Added

  • AuthenticatedClient(token=...) now works out of the box against production: base_url defaults to https://api.bytekit.com.
  • Finite default request timeout of 120s (httpx.Timeout(120.0)) — requests no longer hang indefinitely. Override with timeout=.
  • errors.UnexpectedStatus now carries optional .code / .message attributes. The AuthenticatedClient.search() wrapper populates them from the documented Error envelope, and safely raises a typed error (never a bare json.JSONDecodeError) on a non-JSON documented-status body such as an HTML 502.

@hunt-labs/bytekit-cli

Command-line client, published on npm. Binary: bytekit.

[Unreleased]

[0.10.1] - 2026-09-24

Changed

  • A /v1/search provider failure now arrives at HTTP 503 instead of 502 (#4891): Cloudflare replaced the API's 502 body with its own plain-text page, so the search_provider_error code never reached the CLI. The body is unchanged and the CLI already reads it by shape, not status; only a source comment changed in this package.

[0.10.0] - 2026-09-14

Added

  • --solve-challenge on scrape create (#4449). It sends solve_challenge: true, which asks the scrape to attempt to clear a Cloudflare challenge; success is not guaranteed. It is billed as a normal browser render. Without the flag the request body is unchanged and carries no solve_challenge key.

[0.9.4] - 2026-09-10

Added

  • --cursor, --limit and --status on the three bulk recovery commands (#3717): bulk screenshots list <id>, scrape bulk get <id> and fetch bulk get <id>. They forward the cursor pagination the API already exposed, so a job whose webhook manifest came back incomplete can be enumerated page by page from the CLI. --limit is bounded 1-500 and --status accepts only the documented result statuses; both are rejected locally before a request is made. No new command and no --all flag — the existing commands gained options, nothing else moved.

[0.9.3] - 2026-09-03

Changed

  • Breaking (API): --format now takes raw | markdown | links | images (#3775). raw is the renamed unprocessed-source format, and the cleaned-article-HTML format is gone; the server rejects both retired values rather than translating them, so --format raw --format markdown replaces the old spelling. --raw/-o accept the two textual formats, markdown and raw.

Fixed

  • A failed request's response headers now survive on the thrown ByteKitError (#3784). The CLI's own copy of the SDK error parser dropped the response after reading its body, so Retry-After and the X-Quota-* trio were unreadable on an error even though the gateway sent them. The SDK has passed them into the fifth constructor argument since #2957; this copy now matches, and the parity suite pins the two parsers against each other on a shared header corpus.

[0.9.2] - 2026-08-31

Changed

  • Corrected the detailValueText doc comment in src/output.ts, which claimed the object arm of the details renderer carried the SAME depth bound as the array arm through detailJson (#3691). It carries the same constant (DETAIL_MAX_DEPTH), not the same counter: detailJson takes no depth argument and starts a fresh budget from its own replacer chain. The comment now records the two consequences — the arms agree level-for-level only on homogeneous chains, and a mixed array/object chain can render up to 2 × DETAIL_MAX_DEPTH levels — and why threading depth into detailJson is not the fix: it would elide shapes that render today, breaking the byte-identical inverse-case contract #3647 shipped under. No consumer-visible change — this release touches one comment, and every details shape renders byte-identically to 0.9.1.

[0.9.1] - 2026-08-27

Fixed

  • A rejected request now names the field that was rejected. When the API answers 422 validation_error, its envelope carries details.fieldErrors — the field the server refused and what it will accept. The CLI dropped it twice over: its own transport parser never read error.details off the envelope (the SDK has read it since @hunt-labs/[email protected]), and handleError rendered only the message, code and status. The result was Error: Invalid request body. (validation_error, status: 422) and nothing else, in human and --json modes alike, leaving curl as the only way to find out which field was wrong (#3538).

    $ bytekit scrape create --url https://example.com --country zzzz
    Error: Invalid request body. (validation_error, status: 422)
    Details:
      fieldErrors.country: String must contain exactly 2 character(s)

    The block is rendered on stderr, below the existing one-liner, and is bounded — at most ten lines of at most 200 characters, with a … and <n> more line when there are further entries — so an oversized or deeply nested details can never become a multi-KB dump in a terminal. Nothing else changed: the Error: prefix, the (<code>, status: <n>) suffix, exit code 1 and the silence of stdout are all as before, and an error carrying no details (or a details that is not a JSON object) prints exactly the single line it printed in 0.9.0. --json stdout stays byte-for-byte pipe-clean, because every byte of this goes to stderr.

    The bound covers depth as well as size: the renderer descends at most eight array levels into a details value and elides anything deeper as …. The line and width caps are applied after that walk has finished, so on their own they could not stop a details payload nested thousands of arrays deep — a body JSON.parse accepts, so a server can send one — from overflowing the stack and printing a Node RangeError trace in place of the one-line error (#3633). Every details shape the API sends renders byte-identically: nothing real is more than one array level deep.

    The same bound now covers the object arm of that walk. A details value that is a JSON object was serialized with a bare JSON.stringify, which is itself natively recursive and overflowed the stack on its own: flattening the key by one level bounds WHICH values are rendered, not how deep each one is serialized, so a deeply object-nested payload still printed a RangeError trace (#3647). Objects nested more than eight levels below the flattened key now elide as …, exactly as over-deep arrays do. Everything shallower renders byte-identically, both live shapes included.

[0.9.0] - 2026-08-20

Removed

  • BREAKING: --cookies and --headers are gone from scrape create and screenshots create (#3242). The API no longer accepts caller-supplied cookies or request headers on any capture endpoint, so passing either flag is now an unknown-option error before any request is issued, and a body carrying those keys is rejected by the server with 400 validation_error. There is no deprecation window, feature flag, or compatibility header.

    # before
    bytekit scrape create --url https://example.com --cookies '[{"name":"session","value":"abc"}]'
    # after — the flag does not exist
    bytekit scrape create --url https://example.com

    --webhook-headers on monitors is a different flag and is unaffected.

[0.8.1] - 2026-08-13

Fixed

  • bytekit search explains a malformed 200 instead of leaking a TypeError. A server answering 200 {"ok":true} — a proxy, a stub, a future envelope change — drove the human-mode renderer into Error: response.results is not iterable, which named an internal expression rather than the field the body was missing. It now reports, on one line and with no stack trace, search: the response carries no "results" array — the server returned an unexpected 200 body. Re-run with --json to inspect it. and exits 1. A results value that is present but not an array is caught by the same gate: a string was previously iterated CHARACTER BY CHARACTER, printing one garbled row per character and exiting 0. Unchanged: an empty results: [] is still a successful zero-hit search that prints nothing and exits 0; valid responses render byte-identically; and --json never reaches the renderer, so it keeps dumping the raw body at exit 0 — which is how you inspect an unexpected body.

Added

  • screenshots … -o <file> now says so when the filename contradicts the captured format. --format is omitted far more often than not and the server default is jpeg, so bytekit screenshots create --url … -o shot.png wrote JPEG bytes into a .png-named file silently, at exit 0; screenshots get has no --format flag at all, so it could not even be told what it was downloading. Both now print one advisory line to stderr — Note: the screenshot was captured as jpeg, but shot.png names .png. The file holds jpeg bytes — pass --format png to capture it as png. The captured format is read from the artifact URL's own path, which is the only place the CLI can observe it. Advisory, not enforcement: the bytes written and the exit code are unchanged, and the note goes to stderr, so -o - pipelines and stdout redirects stay byte-identical. It stays silent unless both the destination and the artifact name a recognized image format and the two differ — .jpg and .jpeg are the same format, and an extension-less destination, a non-image extension, -o -, or an artifact URL with no usable extension all make no claim to contradict. A not-ready poll is untouched: it still prints its bare id and poll note at exit 2.

Changed

  • The README now documents the whole flag surface, and says which reference wins. Every flag worked and every flag appeared in --help; the README simply framed subsets as exhaustive, so bulk create's --webhook-secret / --defaults / --metadata, scrape bulk create's --custom, sitemap create's --webhook-url / --webhook-secret / --cache-ttl / --compact / --process / --metadata, monitors create and monitors update's --cron (required with --interval-type cron) and --metadata plus the rest of update's surface, fetch create's ten markdown flags and --custom, and screenshots create's --headers / --cookies / --metadata were all undocumented. All of them are documented now, and a note at the top of the README states that --help is the authoritative flag reference. Required flags are marked as such — --file and --webhook-url on all three bulk creators, --url / --interval-type / --webhook-url on monitors create, and --url on scrape create, screenshots create, sitemap create and both fetch verbs — so no reader mistakes one for optional. Note in particular that a bulk job has no poll-only mode: --webhook-url is mandatory there, unlike on sitemap create where omitting it really does mean poll-only. A test asserts the parity in the direction that matters — every flag the parser accepts must appear in the README — so this cannot drift again silently. No command, flag or behavior changed.

  • Alignment-release note, carried forward from this version's original cut. 0.8.0 published to npm at 2026-08-11T21:18:01.304Z. Doc comments in src/client.ts and src/package-version.ts were reworded after the commit that introduced the 0.8.0 version string — they had re-spelled the X-ByteKit-Client header name in prose, which issue #2906's pre-merge gate 1 counts. No command, flag, exit code, or runtime behavior differs from 0.8.0 on account of those comments.

    0.8.1 was cut for that reason because check-release-drift (#2680) is version-gated and anchors on the commit that introduced a version, not on the commit whose tree was actually published. Once 0.8.0 reached npm, any shipped-file change after that anchor reads as post-release drift. Bumping moves the anchor past those comments. 0.8.1 had not yet reached npm when the entries above landed, so they ride this same pending release rather than stranding it as a never-published version.

  • Corrected a false legacy-name claim carried since the rename (#2950). The 0.8.0-era note below said the prior rapidcrawl-cli name "remains installable at its last version but is deprecated." Verified 2026-08-13: rapidcrawl-cli returns npm 404 ("is not in this registry"), and npm deprecate rapidcrawl-cli@'*' fails with the same 404 — there is no published version left for a deprecation notice to attach to, so the claim never held. No registry stub or deprecation was possible to publish (npm has nothing to act on); root CLAUDE.md carried the same false claim and is corrected alongside this entry. Users of the old name should install @hunt-labs/bytekit-cli directly. No command, flag, or runtime behavior changes.

[0.8.0] - 2026-08-11

Added

  • Every outgoing request now carries X-ByteKit-Client: cli/<version>. The CLI identifies itself with the sanctioned client_surface marker (issue #2906, slice 3 of #2701), so bytekit usage is distinguishable in ByteKit's own analytics from both raw API calls (api_direct) and programmatic @hunt-labs/bytekit-sdk usage (sdk_ts) instead of being counted as one of them. cli is a member of the closed client_surface vocabulary declared in issue #2904; the version affix is this package's own version, resolved from package.json at runtime — the same value bytekit --version prints.

    Set at exactly ONE place — the header literal inside the private transport core all three CLI transports funnel through — so it rides on every command, on the JSON path, on the raw /v1/fetch path, and on the 429 retry attempt alike. Nothing else about a request changed: same Authorization and Content-Type, same body bytes, same per-attempt abort budget, same single-retry 429 behavior, same error taxonomy, same exit codes and output shapes. The marker carries no user data and cannot be set, overridden or read from the command line — --headers is a scrape option that travels in the request body to the captured site, and never touches the CLI's own request headers.

    Minor rather than patch because the wire format of every request changes, even though no documented command behavior does. A gateway that does not yet recognize the header ignores it.

[0.7.1] - 2026-08-10

Changed

  • Version-alignment release after the 2026-08-10 publish rescue: the published 0.7.0 artifact was built from the main promotion HEAD, which already contained every change documented under the 0.7.0 heading — but that commit postdates the one that introduced the 0.7.0 version string, so the check-release-drift gate (#2680) correctly reported shipped drift on every subsequent CI run (#2829). 0.7.1 realigns manifest ↔ registry ↔ version-introducing commit. No consumer-visible change relative to the published 0.7.0.

[0.7.0] - 2026-08-07

Breaking

  • A not-ready screenshots get -o / screenshots create -o now exits 2 and prints the bare ss_ id, instead of exiting 1 with not ready yet: screenshot is pending. The two poll commands disagreed: scrape get --raw on a still-queued scrape already wrote the bare sc_ id to stdout and exited 2 — a distinct, scriptable "accepted but not ready" signal — while the screenshots side reported the same state as a failure and threw the handle away, so a polling script could branch on "queued" for one command and not the other. The screenshots branch now mirrors scrape exactly: bare ss_ id on stdout, a poll note on stderr, exit 2, and (as before) no file written. A script that treated any non-zero exit as "this screenshot failed" now sees 2 for a job that is merely still running — branch on the code (2 = poll again, 1 = failed) rather than on non-zero. Three states deliberately do not move: a terminally failed screenshot stays exit 1 with its error surfaced (it will never complete, so a "poll again" signal would loop a script forever), an envelope carrying no ss_ id stays exit 1 (there is no handle to hand back), and --json keeps the envelope at exit 0.
  • --json on a fetch command now emits X-Fetch-ID and X-Fetch-URL, not X-Fetch-Id and X-Fetch-Url. The metadata keys were derived by a generic segment-wise Title-Caser, which lowercases every non-initial character of a segment — so the two acronym-bearing headers came out in a spelling the API reference never uses, and a jq '."X-Fetch-ID"' filter written from the docs returned null. All 16 X-Fetch-* keys now match the spec names byte-for-byte. The values are unchanged, the non-X-Fetch-* headers stay filtered out, and every other key (X-Fetch-Waf-Vendor, X-Fetch-Duration-Ms, X-Fetch-Fast-Path, …) is spelled exactly as before. A consumer reading .["X-Fetch-Id"] must switch to .["X-Fetch-ID"].

Fixed

  • A screenshots get -o / screenshots create -o poll only says "poll again" for a pending or processing screenshot. The not-ready predicate decided from the presence of an image_url, an id and a nested error alone — status was read only to phrase the message — so a completed screenshot whose image_url never materialized, a status outside the closed enum, and the flat error_code/error_message failed shape all received exit 2 and "no image artifact is available yet. Poll it with: …" for a job that would never complete. Those states now exit 1: a terminal failure prints screenshot failed: …, and any other artifact-less state prints no image artifact available: screenshot is <status>. Genuine pending / processing polls keep the bare ss_ id, the poll note and exit 2 exactly as before — including a poll whose envelope carries an error_code from an attempt that failed and is being retried, since the status, not the error, decides whether polling can still help. --json still prints the envelope at exit 0.
  • A terminally failed screenshot reported through the spec's flat error_code/error_message fields now reaches stderr. Only the nested {error:{code,message}} envelope was read, so the shape ScreenshotResponse actually declares surfaced neither code nor message. It now prints screenshot failed: <message> (<code>) — the formatting scrape get --raw already uses for a failed scrape. The nested envelope's own message is unchanged.
  • scrape get --raw on a terminally failed scrape now reports why it failed. It printed --raw: response contains no formats — technically true of a ScrapeErrorEnvelope, and the one diagnostic that cannot explain anything — while the envelope's own error.code and error.message never reached the caller. Both now go to stderr (exit 1, unchanged). The failure is reported ahead of the format diagnostics, so an envelope carrying both a failure and an empty formats: {} still says why it failed rather than pointing at a --format choice that cannot help.
  • An empty-formats success now names the warning that explains it. A completed scrape whose persisted artifact could not be rehydrated comes back with formats: {} and an artifact_unavailable warning; --raw rendered that as --raw requires --format when the response contains multiple formats, which is false on both counts. It now names the warning's code and message, or says plainly that the formats map is empty when no warning explains it.
  • Error paths no longer risk losing their own diagnostic. handleError and the missing-API-key path each wrote to stderr and then immediately hard-process.exit(1). Node's stderr is asynchronous when it points at a pipe, and process.exit discards whatever is still queued — so the one line a script greps was the line most at risk of being thrown away, which is exactly the defect this package's flush policy exists to prevent. Both now set process.exitCode and let the process end naturally once stderr has drained. The exit codes and the rendered messages are byte-identical.
  • --headers, --custom and --metadata said must be valid JSON (a object) on a malformed value; the article now agrees — an object / an array.
  • The [0.6.0] link reference, missing from the bottom of this file, is restored.

[0.6.0] - 2026-07-29

Breaking

  • screenshots create --async now exits 2 in human mode instead of 0. The queued 202 is not a completed capture — there is no artifact yet — and every other queued path in the CLI already signalled that with exit code 2 (scrape create --async). A script that relied on exit 0 to mean "the screenshot is ready" was reading a false success; under set -e such a script now stops at the queue step, which is the intended correction. The envelope itself is unchanged and still printed. Two carve-outs: --json keeps exit 0 (the envelope is a complete machine-readable payload — identical to scrape create --async --json), and --async -o <file> also exits 2 while still writing no file (it prints the queued envelope and skips the download, as before). Poll with bytekit screenshots get <id>.
  • A --file per-item key outside a bulk endpoint's own strict schema now fails the whole create with a 422 instead of silently succeeding with that key ignored. In 0.5.x, --file never read anything but url from a JSON line — every other key was dropped client-side before the request was ever built. As of 0.6.0 those keys ride to the server (items on bulk create / scrape bulk create; inside urls on fetch bulk create), and /v1/bulk / /v1/scrape/bulk reject any key their .strict() per-item schema doesn't recognize. A file that worked in 0.5.x purely because its extra keys were quietly discarded may now 422 — see each command's --file help for the keys actually honored on that endpoint.
  • fetch bulk create --file per-item keys outside url/format/country/cache_ttl now fail the whole create with a 422 instead of silently succeeding with that key ignored. /v1/fetch/bulk's per-item override schema (inside urls[]) is now .strict() on the gateway (issue #2648), matching the /v1/bulk / /v1/scrape/bulk posture above. The --file help text for fetch bulk create is now uniform with bulk create / scrape bulk create and no longer enumerates keys — this entry is where the honored set (url, format, country, cache_ttl) is documented instead.

Added

  • Per-item bulk configuration on all three bulk creates (bulk create, scrape bulk create, fetch bulk create). A --file line may now be a JSON object carrying url plus per-item fields instead of only a bare URL; such a file is sent as items on /v1/bulk and /v1/scrape/bulk, and as a heterogeneous urls array on /v1/fetch/bulk (which has no items property). A line whose only key is url still normalizes to a bare URL string, so a file of url-only lines produces a byte-identical request to before. A JSON line with a missing or non-string url is now rejected locally with a usage error (exit 1) instead of being sent to the server as a URL. (See Breaking, above, for the per-item strict-schema behavior change this enables.)
  • --defaults, --webhook-secret and --metadata on all three bulk creates, plus --custom on scrape bulk create only (/v1/bulk and /v1/fetch/bulk have no top-level custom, so the flag would be rejected or stripped there). --defaults is forwarded verbatim; the server owns precedence and unknown-key handling.
  • The full markdown-tuning surface on fetch create: --markdown-query, --markdown-links, --markdown-images, --with-links-summary, --with-images-summary, --markdown-compact, --markdown-filter-images, --markdown-include-media, --markdown-include-warnings, --markdown-include-stats, and --custom.
  • Webhook delivery and processing flags on sitemap create: --webhook-url, --webhook-secret, --cache-ttl, --compact, --process, --metadata. Webhook-driven sitemap crawls were previously unreachable from the shell.
  • --headers, --cookies and --metadata on screenshots create.
  • --metadata on monitors create, and --metadata + --url on monitors update. --type remains deliberately absent from update — it is immutable server-side.

[0.5.1] - 2026-07-28 [NEVER PUBLISHED]

Never published to npm. This version was cut in-tree but no @hunt-labs/[email protected] tarball exists — npm i @hunt-labs/[email protected] fails. Everything below shipped to users in 0.6.0, which is the first published release containing it. Reconciled by issue #2680; the release-drift CI gate now blocks a new phantom heading from being introduced.

Fixed

  • Corrects the scope of the 0.5.0 "stalled body" fix claim. screenshots create/get -o downloads the image_url artifact bytes with a bare globalThis.fetch call (lib/download-artifact.ts) that sat entirely outside the transport seam sendWithRetry owns — no abort signal, no connection-error taxonomy. The 0.5.0 entry below ("The per-attempt request timeout now covers the response body read... a server that answered with headers and then stalled its body... hung the CLI forever", #2546) reads as if it covered every response body; it only ever covered API-call bodies read through the JSON/raw transports. A stalled artifact host (headers, then a body that never completes) still hung the CLI indefinitely until now. The download path now shares the same per-attempt abort budget and ByteKitConnectionError timeout taxonomy as the API transport (createAttemptBudget / toConnectionError, exported from client.ts — no second timeout implementation), with the same 120 s per-attempt default and no retry (artifact GETs are not the API's 429 domain). Partial-file policy: the response is buffered fully in memory before any write, so an aborted download never leaves a partial/truncated file at the -o destination. User-visible consequence: a legitimately slow (but not stalled) large artifact download now hard-fails after 120 s instead of waiting indefinitely, and there is currently no CLI flag to raise that budget for a single command. (#2587)
  • While in the file: the 429-retry path left its (unread) response body undrained before sleeping and retrying — undici cannot reclaim the underlying socket for connection-pool reuse until a body is fully read. The body is now drained with response.text() on that path, bounded by its own short (2 s) cap independent of the 120 s per-attempt budget, so a 429 whose body stalls can no longer delay the retry by anywhere near the full attempt budget (or, with the budget disabled, forever). Correction to an intermediate Unreleased draft of this same fix: that draft's cap only bounded how long the retry waited — it resolved on schedule but never touched the abandoned response.text() call itself, which kept reading against a live socket indefinitely. Because every CLI success path exits via a soft process.exitCode (#2544) rather than process.exit(), that orphaned read kept the whole process alive forever — strictly worse than having no drain at all, and on the exact class of hang this ticket exists to eliminate. The cap now also aborts the request when it fires, so the abandoned read is actually torn down, not merely ignored; this holds even when the per-attempt timeout itself is disabled (timeoutMs <= 0). User-visible consequence: a 429 with a stalled body still costs the retry at most ~2 s, exactly as before, but the CLI process now reliably exits afterward instead of hanging forever regardless of that visible latency. The same discard-drain on the artifact-download's non-2xx path is removed outright rather than capped: that path throws straight into process.exit(1), so there is no subsequent request in the process for a reclaimed socket to serve, and the drain there bought nothing. (#2587)

[0.5.0] - 2026-07-24

Note

  • The 0.3.1 and 0.4.0 versions below document real, tagged changes but were never published to npm — the registry's version history jumps straight from 0.3.0 to 0.5.0. Don't go looking for those tarballs; the changes they describe ship for the first time in this release.

Added

  • Surface parity round 2 (#2545): every x-sdk-scope: v0.1 endpoint the SDK exposes now has a CLI command, and the high-value request fields are reachable by flags. Each new flag follows the omit-when-absent rule — absent → the key is left out of the request, so a server default is never materialized.
    • New command groups: usage get | daily | by-endpoint (billing-period usage reporting), webhook-deliveries list | retry <id> (inspect and requeue deliveries), and bulk cancel <id> (stop an in-flight bulk job).
    • Monitor pause/resume: monitors update --status paused / --status active. monitors create/update also gain --change-threshold (float), --notify-on, --webhook-secret, and the JSON-object flags --webhook-headers, --options, --scrape-options. --type is create-only (a monitor's type is immutable).
    • scrape create multi-format --format: repeatable AND comma-separated — --format raw_html --format markdown and --format raw_html,markdown both send formats: ["raw_html","markdown"]. Raw output (--raw/-o) requires exactly one format. New scrape flags: --cache-ttl, --markdown-query, --cookies, --headers, --delay-ms, --events, --token-encoding, --custom, --remove-base64-images / --no-remove-base64-images (tri-state), and the --markdown-* / --with-*-summary family.
    • scrape get --raw / -o / --format: retrieve a completed queued scrape's body from the shell (a still-queued job prints the id and exits 2).
    • screenshots create: --block-ads/--no-block-ads, --block-cookie-banners/--no-block-cookie-banners, --scroll/--no-scroll (tri-state pairs), --wait-for-selector, --delay-ms, --dark-mode, --device-scale-factor, --include-html, --country, --language.
    • fetch create / fetch get: --country, --timeout-ms, --cache-ttl (both verbs); --markdown-mode (create only).
    • sitemap create: --strategy, --max-depth, --max-urls.

Fixed

  • --json no longer emits the literal text undefined — which is not valid JSON — for a response with no body. JSON.stringify(undefined) returns the value undefined, so bytekit monitors delete <id> --json | jq . failed with a parse error on every delete, and --json -o <file> on the same response crashed with ERR_INVALID_ARG_TYPE instead of writing anything. Every --json code path now serializes through one helper. (#2544)

  • scrape create --async --raw (and the bare -o form) no longer risks losing the sc_ id it writes. The id went to stdout immediately followed by a hard process.exit(2); Node documents process.stdout as asynchronous when it points at a pipe, so anything still queued at that moment is discarded — exactly the sc_id=$(bytekit scrape create … --async --raw) capture the flag exists for. The queued signal is now a soft process.exitCode, so the process ends naturally once stdout has drained. The exit code is unchanged. (#2544)

  • The per-attempt request timeout now covers the response body read, not just the time to first byte of headers. The abort budget was armed for fetch and cleared the moment headers arrived, after which reading the body was unbounded — so a server that answered with headers and then stalled its body (including a non-2xx whose error body never finished) hung the CLI forever. The budget now spans headers and body on both transports (request and the /v1/fetch raw path), and a mid-body abort surfaces as the usual ByteKitConnectionError with code: 'timeout'.

    The default is unchanged at 120 s, and the budget is still per attempt — the single 429 retry and its backoff sleep still get a fresh full budget. One consequence worth planning for: a legitimately large body (a multi-MB /v1/fetch document) now counts against that same 120 s, where previously only the wait for headers did. (#2546)

Documented

  • Output contract: with --json, a 204 No Content response (e.g. monitors delete) renders as JSON null. Human mode still prints nothing for the same response. (#2544)
  • Exit codes now have their own README section. In particular exit code 2 — "accepted but still queued", emitted by scrape create --async in raw and human modes — was previously documented nowhere despite being observable since 0.3.0. It is deliberately neither 0 (a content-body success) nor 1 (an error). (#2544)
  • Corrected two false README claims about scrape output: a bare -o implies --raw and writes the requested format's body (it does not write "the same text rendering the terminal would show"), and --async --raw prints the bare sc_ id on stdout with a note on stderr (it does not print the queued job envelope). (#2544)

[0.4.0] - 2026-07-23 [NEVER PUBLISHED]

Never published to npm. This version was cut in-tree but no @hunt-labs/[email protected] tarball exists — npm i @hunt-labs/[email protected] fails. Everything below shipped to users in 0.5.0, which is the first published release containing it. Reconciled by issue #2680; the release-drift CI gate now blocks a new phantom heading from being introduced.

Removed (BREAKING)

  • The entire bytekit account api-keys command group — create, list, get, update, reveal, and revoke. Every one of these subcommands was structurally unreachable: they target /v1/account/api-keys*, which the API mounts behind a Clerk dashboard session and marks x-internal in the OpenAPI spec. The CLI's only credential is a bearer sk_ API key, which can never satisfy a session verifier, so all six returned 401 unauthorized in every published version. Running one now exits 1 with error: unknown command 'api-keys'.

    Key management was deliberately NOT moved onto bearer auth. The reveal operation returns the live plaintext of any key on the account, and there is no per-key scope or permission model, so accepting API-key auth there would turn any leaked key into a self-perpetuating credential factory.

    Replacement: manage API keys from the dashboard at app.bytekit.com. (#2536)

[0.3.1] - 2026-07-22 [NEVER PUBLISHED]

Never published to npm. This version was cut in-tree but no @hunt-labs/[email protected] tarball exists — npm i @hunt-labs/[email protected] fails. Everything below shipped to users in 0.5.0, which is the first published release containing it (0.4.0 was never published either — see its heading above). Reconciled by issue #2680; the release-drift CI gate now blocks a new phantom heading from being introduced.

Fixed

  • bytekit fetch get / fetch create print the fetched content again. 0.3.0's claim that fetch "prints the raw body by default, or the X-Fetch-* metadata as JSON with --json" was not true of the shipped artifact: the CLI replaced only client.request, but SDK 0.3.0 moved /v1/fetch onto client.requestRaw, so the CLI's raw-body branch was unreachable and every fetch fell through to the generic key/value printer — emitting body: …, a JSON map of every byte under bytes:, and headers: {}. The CLI now overrides BOTH transport seams and renders the SDK's real FetchResult. Binary content is byte-exact in every mode. (#2535)
  • Fetch commands honor the CLI's single 429 retry again. Riding the un-overridden SDK transport also bypassed the retry, the per-attempt 120 s timeout, and the connection-error taxonomy; all three now apply to fetch exactly as they do to the JSON commands. (#2535)

Added

  • -o, --output <file> on fetch get / fetch create: writes the fetched body bytes to a file, byte-for-byte (with --json, writes the X-Fetch-* metadata object instead). (#2535)

[0.3.0] - 2026-07-21

Removed (BREAKING)

  • The legacy RAPIDCRAWL_API_KEY and RAPIDCRAWL_BASE_URL environment-variable fallbacks (kept as a one-minor-version bridge in #2421) are removed. Use BYTEKIT_API_KEY and BYTEKIT_BASE_URL (or the --key / --base-url flags). (#2473)

Fixed

  • bytekit fetch get / fetch create no longer crash on non-JSON bodies. The transport JSON-parsed every 2xx response, but /v1/fetch returns the fetched content AS the body. The fetch command now prints the raw body by default, or the X-Fetch-* metadata as JSON with --json; fetch bulk and every other command still parse JSON. (#2467)
  • 204 No Content / empty-body responses (e.g. monitors delete, account api-keys revoke) no longer throw a JSON-parse error, including after a 429 retry. (#2467, #2468)
  • search --type images no longer prints an undefined snippet line for results that omit the field. (#2479)

Added

  • scrape create accepts the full ScrapeRequest surface: --country, --timeout-ms, --markdown-mode, --async, --webhook-url, --token-budget, --clean-markdown, --include-tags, --exclude-tags, --mobile. Each flag is omitted from the request when absent. (#2480)
  • screenshots create accepts --device, --viewport WxH, --format, --quality, --wait-until, --clip x,y,w,h, and --async; -o <file> on screenshots create/get downloads the captured image bytes (-o - for stdout). -o cannot be combined with --json on screenshots. (#2480)
  • monitors list accepts --limit, --cursor, and --status query filters. (#2480)
  • account api-keys reveal <ak_id>: prints the key masked by default (sk_live_…2345); pass --show for the full plaintext. (#2480)
  • Per-attempt transport timeout (120s, mirroring the SDK default): a hung request now fails with code: 'timeout' instead of hanging forever. Network failures are mapped onto the SDK's connection-error taxonomy (timeout / aborted / connection), and error output renders the API or connection error code. (#2479)
  • A queued async job (HTTP 202) under --raw prints the job envelope with its sc_ id instead of an empty body. (#2479)

Changed (BREAKING)

  • fetch --format is validated locally against the spec union 'markdown' | 'html'; off-spec values are rejected with a clear error before a request is sent (previously forwarded as-is). (#2475)

[0.2.2] - 2026-07-16

Changed

  • Re-release of 0.2.1. The 0.2.1 artifacts never reached npm — the prod publish's publish-sdk step failed on a CI test-isolation flake (cjs-require.test.ts rebuilding the shared dist), which is now fixed. No consumer-facing changes beyond 0.2.1.

[0.2.1] - 2026-07-16 [NEVER PUBLISHED]

Never published to npm. This version was cut in-tree but no @hunt-labs/[email protected] tarball exists — npm i @hunt-labs/[email protected] fails. Everything below shipped to users in 0.2.2, which is the first published release containing it. Reconciled by issue #2680; the release-drift CI gate now blocks a new phantom heading from being introduced.

Changed

  • BYTEKIT_API_KEY / BYTEKIT_BASE_URL are now the primary environment variables (precedence: flag > BYTEKIT_* > legacy RAPIDCRAWL_*). RAPIDCRAWL_API_KEY / RAPIDCRAWL_BASE_URL continue to work as silent fallbacks. --help text, the missing-key error, and the README now reference BYTEKIT_*. (#2421)

Removed

  • The recordings command (the endpoint is permanently 410 deactivated). (#2421)

[0.2.0] - 2026-07-15

Additive, lower-risk release — no breaking changes to existing commands.

Added

  • scrape create --raw / -o, --output <file>: write the raw response body to stdout or a file (#2425).
  • Local search --limit validation via a commander argParser, so an invalid limit is rejected before a request is sent (#2425).
  • One automatic retry on HTTP 429, honoring a clamped Retry-After header, in the CLI transport (#2425).

Fixed

  • Entrypoint guard canonicalized so the bytekit bin runs correctly when invoked through an npm/npx symlink (#2418).

Changed

  • Package renamed to @hunt-labs/bytekit-cli (the @bytekit npm scope was unobtainable); the bin name stays bytekit. The prior rapidcrawl-cli name remains installable at its last version but is deprecated. Corrected in 0.8.1 (see below): that claim does not hold — rapidcrawl-cli 404s on npm and there is no published version left to deprecate.
  • Consumer call sites propagate the SDK's narrowed, spec-derived request types (#2426).
  • Publish-ready packaging metadata: dropped main, unpinned the SDK dependency to workspace:^, added engines, keywords, homepage, bugs, README.md, LICENSE, and an MIT license field.

[0.1.0] - 2026-07-13

Added

  • Initial release of the ByteKit CLI, wrapping the ByteKit SDK as shell subcommands, published under the original rapidcrawl-cli name.

@hunt-labs/bytekit-mcp

MCP server, published on npm.

[Unreleased]

[0.3.17] - 2026-10-02

Changed

  • The bundled documentation snapshot's sdk/python/client page now opens with a runnable install, client, create_scrape.sync and read-the-markdown example before the generated class reference (#5068). The get_doc and search_docs tools serve the new text; no tool, schema or transport change.
  • The bundled documentation snapshot's API reference pages now list the keys of nested object fields as dotted paths under their parent, for example metadata.title and metadata.ogSiteName on the scrape response, instead of naming only the top-level fields (#5069). The get_doc and search_docs tools serve the new text; no tool, schema or transport change.
  • The bundled documentation snapshot's introduction page now carries a Python SDK example (AuthenticatedClient, create_scrape.sync, ScrapeRequest) after the TypeScript one in its "Integrate in 60 seconds" block (#5070). The get_doc and search_docs tools serve the new text; no tool, schema or transport change.
  • The bundled documentation snapshot's introduction page now requests the markdown format in its TypeScript and Python examples and reads the result from the success envelope (result.metadata.title, result.formats.markdown), so a copied snippet shows the response shape (#5072). The get_doc and search_docs tools serve the new text; no tool, schema or transport change.

[0.3.16] - 2026-10-01

Changed

  • The bundled documentation snapshot's bulk API reference pages now list total_billed_bytes on the bulk job bodies (#5021). The get_doc and search_docs tools serve the new text; no tool, schema or transport change.
  • The bundled documentation snapshot picks up the Prettier formatting of the docs pages (#5034), which changes the text served for billing/bandwidth, guides/errors, guides/monitors, guides/rate-limits, guides/scraping, guides/webhooks and mcp. No tool, schema or transport change.
  • Refreshed the bundled documentation snapshot after the billing/balance page's one-time starter grant for new pay-as-you-go accounts rose from 50 MB to 100 MB of bandwidth, with the 100-credit grant unchanged (#4959). No behavior change in the package: bundled content only, no tool, schema or transport change.

[0.3.15] - 2026-09-30

Fixed

  • The docs-bundle builder now reads the llms.txt index in its new - [Title](url): description link shape and ignores URLs outside /docs (the new summary block's api.bytekit.com and app.bytekit.com), so regenerating the bundled documentation snapshot after the docs llms.txt reorder (#5007) still yields the same page set. No tool, schema or transport change.

Changed

  • The bundled documentation snapshot now carries the rewritten Introduction page, which states the base URL, the BYTEKIT_API_KEY auth header, the install lines for both SDKs, one curl call and one TypeScript SDK call, so an agent that reads only that page can integrate ByteKit (#5005). The get_doc and search_docs tools serve the new text; no tool, schema or transport change.
  • The bundled documentation snapshot now carries the SDK package names on the Client Libraries, TypeScript SDK, Python SDK and Quickstart pages: install @hunt-labs/bytekit-sdk (npm) / bytekit-sdk (PyPI), and the npm package named bytekit is not ByteKit's (#5004). The get_doc and search_docs tools serve the new text; no tool, schema or transport change.

[0.3.14] - 2026-09-29

Changed

  • Pinned @modelcontextprotocol/sdk to the exact version 1.31.0 instead of the range ^1.12.1, so an npx @hunt-labs/bytekit-mcp install and the hosted /mcp endpoint run one SDK version instead of the npx install floating to whatever upstream last published (#4949). The pin moves only when we release a deliberate bump. Earlier releases keep the caret range in their published manifest. No tool, schema or transport change.
  • Refreshed the bundled documentation snapshot after the top-ups page (billing/top-ups) gained a section on bandwidth owed: an overage that no allowance, USD balance or starter grant can cover is recorded as bandwidth you owe, a top-up, renewal or plan upgrade pays it off first, and new billable requests are refused with 402 quota_exceeded while any is owed (#4960). No behavior change: bundled content only, no tool, schema or transport change.

[0.3.13] - 2026-09-24

Changed

  • Refreshed the bundled documentation snapshot after a documentation-wording correction to the POST /v1/schema billing description, which now says the page bandwidth is billed at its upstream wire bytes with no endpoint factor instead of "exactly like /v1/scrape" (#4865). No behavior change: bundled content only, no tool, schema or transport change.

[0.3.12] - 2026-09-23

Changed

  • Refreshed the bundled documentation snapshot after GET /v1/account's AccountSessionResponse gained the required boolean onboarding_completed (#4780). Additive, patch-level: bundled content only, no tool, schema or transport change.

[0.3.11] - 2026-09-21

Changed

  • Refreshed the bundled documentation snapshot after the errors guide named /v1/fetch's 422 blocked and upstream_error codes and the getfetch and postfetch pages declared the optional X-Fetch-ID header on their 422 responses (#4732). Bundled content only: no tool, schema or transport change.

[0.3.10] - 2026-09-21

Changed

  • Refreshed the bundled documentation snapshot after GET /v1/fetch and POST /v1/fetch started answering a fetch that exhausted its retries against the target with 422 (code blocked or upstream_error) instead of 503 internal_error (#4669); the bundled getfetch and postfetch pages name the new condition. Bundled content only: no tool, schema or transport change.

[0.3.9] - 2026-09-18

Changed

  • Refreshed the bundled documentation snapshot after a GET /v1/bulk/{id}/screenshots item started carrying billing: null while its status is not completed (#4582). Bundled content only: no tool, schema or transport change.
  • Refreshed the bundled documentation snapshot after POST /v1/schema started rejecting slow-path rendering options with error code unsupported_option instead of unsupported_url (#4584); the bundled createschemaextraction page names the new code. Bundled content only: no tool, schema or transport change.
  • Dropped the changelog page from the bundled documentation snapshot (#4675): the page concatenated four package CHANGELOG.md files into one entry, making it the single most merge-collided entry in the bundle. search_docs's description no longer describes a bounded window into release history; it now says release history isn't bundled at all and points to the public changelog page instead.

[0.3.8] - 2026-09-17

Changed

  • Refreshed the bundled documentation snapshot after the TypeScript SDK changelog entry for #4563, which records the widened internal GET /v1/bulk status vocabulary (cancelled is now a documented status and an accepted filter). That entry is inside the bundled body of the changelog page, and the page's truncation notice reports its new full size. Bundled content only: no tool, schema or transport change.

[0.3.7] - 2026-09-14

Added

  • scrape_url accepts an optional solve_challenge boolean (#4450): attempt to clear a Cloudflare challenge; success is not guaranteed. Billed as a normal browser render. The parameter is forwarded to POST /v1/scrape only when set; omitting it sends no key.

Changed

  • Refreshed the bundled documentation snapshot after solve_challenge was published on POST /v1/scrape: the API reference page for that operation and the TypeScript SDK scrape page list the new option. This refresh itself changes no tool, schema or transport.
  • The TypeScript SDK changelog entry for #4447 is inside the bundled body of the changelog page.
  • The Python SDK changelog entry for #4448 is past the cut, so a client reading the bundled changelog page does not see it.
  • Refreshed the bundled documentation snapshot after credits_scrape was published on GET /v1/usage (#4452): the API reference page for that operation lists the new field. This refresh itself changes no tool, schema or transport.
  • The TypeScript SDK changelog entry for #4452 is inside the bundled body of the changelog page.
  • The Python SDK changelog entry for #4452 is past the cut, so a client reading the bundled changelog page does not see it.
  • Refreshed the bundled documentation snapshot after empty_extraction was documented on the POST /v1/schema response (#4455). This refresh itself changes no tool, schema or transport.
  • Refreshed the bundled documentation snapshot after GET /v1/logs billing_multiplier was redocumented as the bandwidth factor derived from the settled bytes, in place of the row's persisted credit count (#4469). This refresh itself changes no tool, schema or transport.
  • The TypeScript SDK changelog entry for #4469 is inside the bundled body of the changelog page.
  • Refreshed the bundled documentation snapshot after the POST /v1/schema validated description recorded the explicitly-nullable exception to the all-null miss (#4490). This refresh itself changes no tool, schema or transport.
  • The TypeScript SDK changelog entry for #4490 is inside the bundled body of the changelog page.
  • The TypeScript SDK changelog entry for #4455 is inside the bundled body of the changelog page.
  • Refreshed the bundled documentation snapshot after an inline POST /v1/scrape failure started carrying a persisted sc_ id (#4454): the bundled api/scrape/createscrape and sdk/python/scrape pages no longer say id: null. This refresh itself changes no tool, schema or transport.
  • The TypeScript SDK changelog entry for #4454 is inside the bundled body of the changelog page.
  • The Python SDK changelog entry for #4454 is past the cut, so a client reading the bundled changelog page does not see it.
  • Refreshed the bundled documentation snapshot after the internal GET /v1/logs feed started listing search requests and webhook deliveries (#4460). Only the bundled changelog page moved; this refresh itself changes no tool, schema or transport.
  • The TypeScript SDK changelog entry for #4460 is inside the bundled body of the changelog page.
  • Refreshed the bundled documentation snapshot after the generated sdk/python/scrape reference stopped listing request fields the OpenAPI spec marks x-internal: true (#4515): that page loses its x_internal_scrape_path bullet. The field remains a live internal control in the spec and in the generated model, and the API reference page for POST /v1/scrape still lists it. This refresh itself changes no tool, schema or transport.
  • The MCP server changelog entry for #4515 is past the cut, so a client reading the bundled changelog page does not see it; the only other bundled bytes that move are in that page's truncation notice, which reports the full size of the page it truncates.
  • Refreshed the bundled documentation snapshot after the internal GET /v1/bulk operation became cursor-paginated and lost its 100-job cap (#4508): the bundled api/bulk/listbulk page now documents the { data, next_cursor, has_more } page envelope and the limit / cursor query parameters in place of a bare array, and both that page's title and the api/bulk index card drop the wrong "in-flight" wording — the operation has always returned jobs of every status. This refresh itself changes no tool, schema or transport.
  • The TypeScript SDK changelog entry for #4508 is inside the bundled body of the changelog page.

[0.3.6] - 2026-09-12

Changed

  • Refreshed the bundled documentation snapshot after the Python SDK changelog entry for #4340, which records the split bulk list-item models and the new BulkCompletedWebhook model. The changelog page concatenates four package changelogs and is cut at the 16,384-byte per-page bound; the Python SDK entry is past the cut, so the only bundled bytes that move are in the truncation notice, which reports the FULL size of the page it truncates. Bundled content only: no tool, schema or transport change.
  • Refreshed the bundled documentation snapshot after the audience-boundary audit of the published docs corpus (#4381). The audit rewrote passages on the billing, scraping, monitors, MCP-tool and command-line pages, and on the generated TypeScript client reference, that described how ByteKit is implemented — internal storage tables and how stored values are written, internal service ownership, operator-only webhook-ingress routes, ByteKit's own staging host, repository provenance and private issue-tracker numbers — rather than what a reader should do or expect. Every public rule, value, API name and error code in those passages is retained, response fields named in the OpenAPI spec included; only the implementation detail is gone. No tool, schema or runtime behavior in this package changed.

[0.3.5] - 2026-09-06

Changed

  • Refreshed the bundled documentation snapshot after the TypeScript and Python SDK changelog entries for #4251, which correct GET /v1/usage's topup_balance_bytes description to name the one-time starter grant alongside the purchased top-up lots it already named. The only bundled bytes that move are on the changelog page. That page concatenates four package changelogs and is cut at the same 16,384-byte per-page bound: the TypeScript SDK entry lands near the top of the page, inside the bundled body, while the Python SDK entry is past the cut and is only reachable at https://bytekit.com/docs/changelog. Bundled content only: no tool, schema or transport change.
  • Refreshed the bundled documentation snapshot after the TypeScript and Python SDK changelog entries for #4237, which record that GET /v1/scrape/{id} reports cache: hit without the scrape envelope's cache_age_s field. The only bundled bytes that move are on the changelog page. That page concatenates four package changelogs and is cut at the same 16,384-byte per-page bound: the TypeScript SDK entry lands near the top of the page, inside the bundled body, while the Python SDK entry is past the cut and is only reachable at https://bytekit.com/docs/changelog. Bundled content only: no tool, schema or transport change.
  • Refreshed the bundled documentation snapshot after the TypeScript and Python SDK changelog entries for #4203, which record GET /v1/scrape/{id} as a third delivery path for the scrape envelope's cache field. The only bundled bytes that move are on the changelog page. That page concatenates four package changelogs and is cut at the same 16,384-byte per-page bound: the TypeScript SDK entry lands near the top of the page, inside the bundled body, while the Python SDK entry is past the cut and is only reachable at https://bytekit.com/docs/changelog. Bundled content only: no tool, schema or transport change.
  • Refreshed the bundled documentation snapshot after the TypeScript SDK's 0.9.4 changelog entry for #3984. The only bundled bytes that move are on the changelog page. That page concatenates four package changelogs and is cut at the same 16,384-byte per-page bound; the added entry lands near the top of the page, inside the bundled body. Bundled content only: no tool, schema or transport change.
  • Regenerated the bundled documentation corpus so the offline docs describe monitor scrape_options correctly (#4024): the description no longer lists headers among the accepted fields, since POST /v1/monitors never honoured a caller-supplied headers option and now rejects it with the same 400 every other capture endpoint returns for the field. No tool, argument or output shape changed.
  • Refreshed the bundled documentation snapshot after the TypeScript SDK's 0.9.4 changelog entry for #4030, which records the third case in which X-Scrape-Proxy-Tier-Label reports no tier: a response replayed from stored bytes after this request's own cache lookup missed, which arrives as X-Scrape-Cache: miss. Bundled content only: no tool, schema or transport change.
  • Refreshed the bundled documentation snapshot after the TypeScript SDK's 0.9.4 changelog entry for #4132, which qualifies that third case as an internal deploy-skew safeguard no API request can reach. Bundled content only: no tool, schema or transport change.

[0.3.4] - 2026-09-04

Changed

  • Regenerated the bundled documentation corpus so the offline docs describe the proxy-tier response headers as they now behave (#3926). X-Scrape-Proxy-Tier-Label reports the tier that actually served a scrape — after any escalation, never the tier first attempted — and is omitted on a cache hit, which replays stored bytes and runs no tier. No tool, argument or output shape changed.

[0.3.3] - 2026-09-03

Changed

  • Breaking (API): scrape_url's formats enum is now raw | markdown | links | images (#3775). The alias layer that accepted the two retired raw-HTML spellings from an agent and translated them onto the wire is gone with the spellings themselves — there is one vocabulary now, so the tool neither translates nor tolerates. A secondary raw format is still rendered under its own --- format: raw --- delimiter; the cleaned-article-HTML format no longer exists.

Fixed

  • A 402 now reaches the agent with the server's OWN error code and message (#3784). The client carried a hardcoded 402 branch that rewrote every payment-required response to quota_exceeded with client-authored bandwidth wording, so an impaired or expired subscription — and an overage spending cap — all read as "Bandwidth quota exhausted". The branch is gone; the generic envelope path reports what the gateway sent, so the four billing codes the API documents are the four codes an agent sees.

Changed

  • Refreshed the bundled documentation snapshot for the two subscription-entitlement 402 codes now documented in the errors guide (#3784).

  • Refreshed the bundled documentation snapshot after the TypeScript and Python SDK changelog entries for #3777, plus this entry itself. The changelog page's truncation notice reports the full page size: 153564 becomes 155329 bytes (+1765, the combined length of the three new entries). The TypeScript SDK's own changelog section alone already exceeds the page's 16,384-byte bundled cut (the Python SDK, CLI and MCP sections that follow it were never in the bundle, before or after this change), so only the TypeScript SDK's new 0.9.2 entry is actually visible — it lands near the top of that section, well inside the cut. Both the Python SDK entry and this entry are further down the page than the cut has ever reached. Pushing everything else in the page down by 1765 bytes moves the cut point 1765 bytes earlier within the TypeScript SDK section's own older history, so slightly less of an older entry is visible than before. The page count stays at 108 and no tool, argument, or response shape changes.

[0.3.2] - 2026-08-31

Changed

  • Refreshed the bundled documentation snapshot after the CLI's new changelog entry (#3691). The only bundled byte that moves is the changelog page's truncation notice, which reports the full page size: 149708 becomes 151210 bytes. That page concatenates four package changelogs and is bound to 16,384 bytes (#3233), so the CLI section itself is still far past the cut and is not in the bundle — the notice's byte count is the whole delta. The page count stays at 108 and no tool, argument, or response shape changes.

[0.3.1] - 2026-08-25

Changed

  • The bundled docs now carry a landing page for every documentation section (#3473). /docs/billing, /docs/guides, /docs/libraries and /docs/sdk had no page of their own and returned 404, so list_docs never offered them and get_doc could not fetch them. Each is now an overview page linking that section's own pages, and the bundle grows from 103 to 107 entries. No tool, argument, or response shape changes.

  • Refreshed the bundled documentation snapshot (#3459). The bundled changelog page is the docs site's changelog, which concatenates four package changelogs in order — @hunt-labs/bytekit-sdk, the Python bytekit-sdk, the CLI, then this package — and every bundled page is bound to 16,384 bytes (#3233), so that page carries the newest TypeScript SDK entries down to the cut and nothing after it. What the refresh actually brings offline is that section's 0.9.1 entry, with the corrected topup_balance_bytes description: the value is the remaining purchased top-up bytes, already net of any overage consumed against them. The Python SDK, CLI and MCP sections are not in the bundle at all — the Python section alone begins roughly 43 KB into the page, far past the bound — so their release notes are not reachable through get_doc or search_docs, and full-text search never sees them. Read the whole changelog at https://bytekit.com/docs/changelog. Bundled content only: no tool, schema or transport change.

[0.3.0] - 2026-08-20

Removed

  • BREAKING: the bundled docs no longer describe cookies or headers as capture request fields (#3242). Those fields were removed from /v1/scrape, /v1/scrape/bulk, /v1/screenshots, /v1/recordings and /v1/schema; a request carrying either is rejected with 400 validation_error. Tools that composed request bodies from the bundle will no longer see the fields offered.

Changed

  • get_doc now returns a bounded body, and says when it cut one (#3233). A page above 16,384 bytes comes back cut short, ending in a notice that names the limit, the page's full size, and the docs URL that serves all of it — a caller can never mistake a partial page for a complete one. Every page under the limit is returned byte for byte, in the same { title, section, body } envelope as before, so nothing else about the tool changes. The cap is advertised in the tool's own description.
  • The bundled offline docs no longer ship an unbounded page (#3233). The harvest step applies the same 16,384-byte per-page bound. The generated changelog page was 46.5% of the whole bundle on its own; it is still a bundled page, so list_docs lists it and search_docs ranks it, but full-text search now reaches only the releases above the cut. The docs site keeps the whole changelog.

[0.2.8] - 2026-08-19

Changed

  • The generated scrape contract's billing factor vocabulary gained phase1_render (#3141). BillingFactorName is a closed union copied verbatim into src/scrape-contract.generated.ts, and screenshots and recordings bill on a basis the byte-scrape factors cannot name — Phase 1 proxy wire bytes times a 1.5 render-overhead multiplier. A capture's billing node carries that one factor and none of the byte-scrape names; the two vocabularies never mix on a single node. This release is a regeneration of the checked-in contract only, with no behavior change in the server itself: the MCP tools do not call the capture endpoints whose responses now carry the node.

[0.2.7] - 2026-08-18

Changed

  • The generated scrape contract now carries the billing node (#3139). ScrapeSuccessEnvelope gained an optional billing object with raw_bytes, an ordered factors list of name/value/reason entries, multiplier and billed_bytes. The factor list explains exactly how a scrape's charge was composed — endpoint factor, cache-hit discount, clean-markdown surcharge — and multiplies out to the X-Billing-Multiplier on the same response. The existing flat billing_multiplier / billed_bytes pair is unchanged and still emitted beside it; this release is a regeneration of the checked-in contract only, with no behavior change in the server itself.

[0.2.6] - 2026-08-12

Changed

  • scrape_url now sends the public raw_html format spelling on the wire (#2959). It previously translated the raw_html an agent asks for into the gateway's internal rawHtml before calling the API. That request only succeeded because the gateway happens to accept both spellings — nothing in the published API contract promised it, so this server depended on an undocumented leniency it could not see. It now sends what the spec documents.

    No change to what the tool accepts: raw_html and the legacy rawHtml are both still valid formats values, and results are rendered identically. The bundled libraries/typescript doc page also gained the search, usage and webhooks resources it had been omitting.

Added

  • The published docs bundle now includes the five billing guides linked from the pricing page: balance, bandwidth, credits, spending, and top-ups. The guides describe the source-backed charging, expiry, rollover, and quota behavior. (#2960)

Fixed

  • A validation failure from the API now reaches the agent as validation_error: <message> from every tool, instead of a bare Invalid request body. ({"formErrors":[],"fieldErrors":{…}}) with the error code stripped off. scrape_url and screenshot_url are the two tools affected: the client's 422 handling rendered a screenshot's failed-capture body as code: message but passed a validation envelope's message through verbatim, so which spelling an agent saw depended on the HTTP status the endpoint happened to pick — the same vocabulary already rendered validation_error: … on a 400. The field details are unchanged, and so are web_search's render, the concurrency-limit advice, and the 401/402 guidance messages. (#2955)
  • The hosted /mcp endpoint now rejects a missing or invalid API key with a JSON-RPC error body ({"jsonrpc":"2.0","error":{"code":-32000,…},"id":null}, HTTP 401) rather than the REST capture envelope, so an MCP host that parses /mcp responses as JSON-RPC can read the failure it is most likely to hit. Every other transport-level failure on that endpoint already spoke JSON-RPC (405 on GET/DELETE, -32700 on a malformed body, 406 on Accept). The ByteKit code (invalid_api_key) is preserved on JSON-RPC's data member, and the capture endpoints' 401 envelope is untouched. (#2955)
  • The 26 SDK reference pages in the docs bundle (sdk/python/*, sdk/typescript/*) no longer open with their title and description twice, so get_doc and search_docs stop spending tokens on a duplicated preamble and search snippets stop reading doubled. The docs site's LLM markdown composer prepended a synthesized # <title> + description block unconditionally, on top of the generator-authored SDK bodies that already carry their own H1 and intro paragraph; it now yields the title block to a body that leads with an H1. The other 77 bundled pages are byte-identical. sdk/python/client's body heading was also aligned to its page title (# Client Reference → # Client), the one SDK page where the two disagreed. (#2953)
  • The screenshot render's signed-URL note no longer presents the plan's artifact-retention date (expires_at, ~30 days out) as the signed image URL's expiry. The URL is a B2-signed link that dies ~10 minutes after issue (X-Amz-Expires=600) regardless of retention; the note now always states that ~600s signature expiry, so an agent following it no longer stashes a link that reads as good for weeks and is already dead. expires_at itself is untouched — it remains correct for plan retention everywhere else. (#2947)
  • An unknown flag (e.g. --bogus-flag) or the space-form --timeout-ms -5 (which node:util's parseArgs itself rejects as an ambiguous option value, before it ever reaches the existing --timeout-ms validation) no longer crashes with a raw node:internal/util/parse_args stack trace. Both now render the same one-line Error: ... message the binary already used for --help, a missing API key, and the equals-form/non-numeric --timeout-ms cases — parse-layer failures and value-validation failures render through one error path. --help, the missing-key message, and the existing --timeout-ms abc / --timeout-ms=-5 errors are unchanged. (#2952)

[0.2.5] - 2026-08-11

Added

  • Every outbound request now carries X-ByteKit-Client: mcp/<package version>, the client_surface attribution marker defined by the Track A analytics contract (#2904), so requests made by the published @hunt-labs/bytekit-mcp server can be told apart from api_direct and the other published clients. Set once, unconditionally, at the client's single header-construction site — it does not ride along with Authorization and is not affected by the in-process gateway dispatch path (issue #1182). (#2907)

[0.2.4] - 2026-08-10

Changed

  • Version-alignment release after the 2026-08-10 publish rescue: the published 0.2.3 artifact was built from the main promotion HEAD, which already contained every change documented under the 0.2.3 heading — but that commit postdates the one that introduced the 0.2.3 version string, so the check-release-drift gate (#2680) correctly reported shipped drift on every subsequent CI run (#2829). 0.2.4 realigns manifest ↔ registry ↔ version-introducing commit. No consumer-visible change relative to the published 0.2.3.

[0.2.3] - 2026-08-07

Added

  • scrape_url / get_result render a zero-count Links (0): / Images (0): header for a format you requested that came back empty, so "the page has none" is no longer indistinguishable from "the format was not returned". A format you did not request still renders nothing at all. (#2687)
  • The queued-scrape render carries the spec-required events array, and the success render carries content_length — the compressed upstream wire-byte count that is the billing basis for the call, which nothing else in the output surfaced. (#2687)
  • Screenshot renders carry page_title (a spec field the tool description already promised as "metadata") and a note that image_url is a signed link with a short lifetime, using the concrete expires_at instant when the API supplies one. (#2687)
  • A queued scrape now gets the same "Use get_result with this ID to check when it completes." guidance a pending screenshot got in 0.2.2. (#2687)
  • screenshot_url exposes quality, dark_mode, country and language. None of them can reach the queued outcome; quality and dark_mode are additionally cost-neutral (a screenshot bills on the phase-1 fetch's wire bytes, never the artifact size), while country/language change which page the origin serves and so can change that figure. token_budget, clean_markdown and cache_ttl stay unexposed on scrape_url, each for a reason recorded in src/tools/scrape-url.ts. (#2687)
  • New output contract. scrape_url / get_result render the envelope's warnings under a Warnings (N): section, one code: message (element) line per warning, on both the success and the terminal-failure render. The (element) suffix is omitted for the whole-extraction codes that carry no element. Nothing rendered them before, so an artifact_unavailable — a format you requested and were billed for is missing from the response — was indistinguishable from a page that simply had none of it. The block is budgeted inside the same 100KB cap as the content, ahead of the links/images sections, so a long links list cannot crowd a failure signal out. A warnings array that is empty or absent renders nothing, exactly as before. (#2679)

Fixed

  • Compile-time breaking for TypeScript consumers that read SearchResult.snippet or SearchResult.date unguarded. Both were declared required on the exported SearchResult type, but POST /v1/search omits both keys for type: 'images' results — docs/api/openapi.yaml declares required: [position, title, url] and nothing else for a search result. Reading result.snippet.length therefore compiled under strict TypeScript and threw at runtime on any images result. They are now snippet?: string and date?: string | null. Runtime behaviour is unchanged — the server has always passed the API's JSON through verbatim, and no MCP tool output moved; what changes is that tsc now requires a guard (result.snippet?.length, or an if (result.snippet !== undefined) narrowing) at every read site. web/news results still carry both fields and date is still nullable there, so null (the provider reported no date) stays distinguishable from absent (the index carries no such field at all). (#2781)
  • The hosted /mcp transport and a local npx install now render invalid-argument errors identically. They had diverged — hosted returned the raw Zod issue-array JSON where stdio returned the flattened Required at url — and the cause was neither a stale deploy nor a difference in ByteKit source: both transports share one server factory, and the text is produced inside @modelcontextprotocol/sdk. What differed was the SDK version each side resolved. @modelcontextprotocol/sdk 1.30.0 rewrote getParseErrorMessage to format Zod issues as <message> at <path>, where 1.29.0 fell through to ZodError.message (the raw JSON array). The hosted image builds with a frozen lockfile, which still pinned 1.29.0, while the unpinned ^1.12.1 range let an npx consumer resolve 1.30.0. The lockfile is now bumped to 1.30.0 so both sides run the same SDK; the declared range is deliberately left at ^1.12.1 so upstream fixes keep arriving. (#2682)
  • A truncated links or images list is cut between whole entries instead of mid-URL, and each affected section is annotated with [Links truncated: N of M entries shown], so the section header's count stays reconcilable with what was actually returned. The markers are spent from the same 100KB budget, not appended on top of it. (#2687)
  • web_search no longer prints the API error code twice (validation_error: … (validation_error)); the code is appended only when the message does not already carry it. (#2687)
  • scrape_url rejects formats: [] at the schema boundary — matching the API's own minItems: 1 — instead of forwarding it and returning a server 422. Omitting the parameter still yields the ["markdown"] default. (#2687)
  • --timeout-ms with a non-numeric value is rejected with an error instead of being silently ignored, which had left the client running on its default budget while the operator believed one was set. (#2687)
  • Reading bytekit://account issues its two sub-requests serially, matching the get_account tool — concurrent sub-requests could rate-limit a 1-rps plan against itself. (#2687)
  • list_docs with a section outside the documentation taxonomy returns an error naming the valid set, rather than a silent empty list that a correctly-spelled but empty section would also produce. (#2687)
  • The Claude Desktop / mcp-remote bridge config in README.md now authenticates as written. It previously passed "Authorization: Bearer $BYTEKIT_API_KEY" in args; an MCP host launches command/args with no shell, and mcp-remote interpolates only the braced ${VAR} form, so the literal text was sent as the bearer token and the endpoint answered 401. The snippet now uses "Authorization:${AUTH_HEADER}" with a matching env block, and the hosted-HTTP snippet carries the same no-shell-expansion caveat. (#2688)
  • The bundled documentation no longer teaches the deactivated /v1/recordings endpoint or links its removed reference pages. (#2688)
  • A 429 caused by the account's concurrency cap now says so and fails fast, instead of being reported as a rate limit. HTTP 429 carries two conditions (rate_limited and concurrency_limit), and the client read neither the body nor the code: every 429 got one Retry-After retry and then Rate limited. Try again in a moment. with the error code hardcoded to rate_limited. The concurrency variant does not clear on that back-off — the API spec calls its Retry-After "a fixed small back-off" while the slots free only when the in-flight jobs finish — so the retry was spent on a condition a second cannot clear, and the error object named the wrong cause. A concurrency_limit 429 now throws immediately with the real wait scale and errorCode: 'concurrency_limit', at both 429 sites. rate_limited 429s, and any 429 whose body is empty, non-JSON or carries no recognizable code, keep the single honored-header retry and the exact message and code they had. (#2681)
  • A screenshot that fails inside the API's sync hold now surfaces its REAL failure code (blocked, timeout, …) and message. POST /v1/screenshots answers that case with HTTP 422 carrying the flat screenshot resource, not the unified error envelope, and the client parsed neither — so every such failure read as the literal Validation error.. The (Screenshot ID: ss_…) handle line is unchanged, and a genuine request-validation 422 still renders exactly as it did. (#2679)

[0.2.2] - 2026-07-29

Fixed

  • scrape_url / get_result render every returned textual format (markdown/html/raw_html), not just the first present — a multi-format call previously dropped every secondary format silently despite it being billed. Secondary formats are rendered under a --- format: <name> --- delimiter, in canonical order after the unlabeled primary. (#2589)
  • The 100KB output truncation budget now covers the links/images blocks too — they were previously appended AFTER truncation, so a combined output could exceed the budget. The whole assembled payload is now truncated exactly once, with the links/images block reserved its own share of the budget so it is never dropped in full just because the textual formats alone already fill the budget. (#2589)
  • scrape_url / get_result descriptions no longer advertise a queued sc_ outcome that the MCP input schema (url/formats/country) cannot produce — that outcome was only reachable via body.async or fields can-fast-path.ts gates on, none of which this tool's schema exposes. (#2589)
  • web_search now renders credits_used and related_searches, both spec-required fields that were previously dropped silently. (#2591)
  • A sync-failed screenshot's ss_ handle — absent from the 422 error body — is now read from the X-Screenshot-ID response header and included in the error text, so it can still be redeemed via get_result. Degrades gracefully when the header is absent. (#2591)
  • The 402 message on /v1/screenshots no longer says "Bandwidth quota exhausted" and no longer points at get_account, which is rate-limit-fragile immediately after a quota failure. Every other endpoint's 402 wording is unchanged. (#2591) (Correction, #3785: this entry originally explained the change by saying screenshot quotas are credit/count-based. They are not. Screenshots are billed on bandwidth — Phase 1 wire bytes x 1.5 — and the credit figure they report is a legacy weight.)
  • get_result's invalid-id-prefix message no longer truncates the offending id to 3 characters. (#2591)
  • get_result now gives the same "poll with get_result" guidance screenshot_url's queued branch gives when a screenshot is still pending/processing. (#2591)
  • get_account no longer renders a duplicate email (Account: [email protected] ([email protected])) when the account has no name. (#2591)
  • search_docs snippets no longer duplicate the page title when the match falls near the body's own leading # <Title> heading. (#2591)
  • ScrapeErrorEnvelope.id is now modeled as nullable, matching docs/api/openapi.yaml (id is null when a scrape was never queued, e.g. a validation failure) — get_result/scrape_url render ID: — instead of the literal ID: null. (#2591)

Changed

  • get_account now issues its getAccount/getUsage requests serially instead of via Promise.all, so it no longer doubles the concurrent request rate against 1-rps plans. (#2591)
  • Regenerated the bundled src/docs-bundle/docs-bundle.json against the current packages/docs/content tree. Every guarded endpoint's docs page now states its request-body byte ceiling and the request_too_large error code (#2640), and the /v1/bulk pages document defaults.type inheritance plus the 422 on unrecognized keys (#2643). Content only — no tool surface, schema, or runtime behavior changed. (#2640, #2643)

[0.2.1] - 2026-07-24

Fixed

  • The spec'd flat /v1/search 502 error body ({"error":"search_provider_error"}) now surfaces its code in the tool error text instead of collapsing to a generic "ByteKit API error (HTTP 502)". Flat {error: "<string>"} bodies on any status parse the same way; envelope-shaped bodies keep winning when both shapes are present. (#2541)

Added

  • BytekitClient accepts a timeoutMs constructor option — a per-request budget in milliseconds, defaulting to 120000. Set 0 to disable it entirely. Values that are negative or non-finite normalize to disabled rather than throwing. Previously a hung upstream hung a stdio MCP tool call forever. (#2568)
  • The budget covers an entire request() call rather than each await separately: the initial fetch, the 429 Retry-After sleep, the retry fetch, and both the success and error body reads (response.json() in request() and in parseErrorEnvelope). Two of those previously swallowed an abort outright — the Retry-After sleep observed no signal, and a failed error-body read was downgraded to "no envelope", surfacing a misleading status-line message instead of the timeout. Exhausting the budget now produces a BytekitApiError with errorCode: 'timeout', distinct from 'network_error'. (#2568)
  • Note the semantics are client-side only: aborting does not cancel or refund server-side work already in progress, so an account may still be billed for a result the caller never reads. The 120000 default is deliberately well above the server-side screenshot ceiling so it never truncates a legitimate billable operation. (#2568)

[0.2.0] - 2026-07-21

Removed (BREAKING)

  • The legacy RAPIDCRAWL_API_KEY environment-variable fallback (kept as a one-minor-version bridge in #2421) is removed. Use BYTEKIT_API_KEY (or the --api-key flag). (#2473)

Fixed

  • Tool-surface drift against the live API (#2477):
    • screenshot_url's device enum is now desktop / mobile — tablet was 422-rejected by the API.
    • Screenshot results render the spec field names (image_width, image_height, file_size_bytes); the removed format field is gone.
    • get_result renders a failed job's envelope (Status: failed + error.code/error.message) instead of undefined lines, and omits Credits: null for pending rows.
    • get_account no longer renders Account: null when the account has no name.
    • get_doc's description example uses a real bundle key.
  • Server errors now surface diagnostics instead of a generic failure: any non-2xx with a parseable API envelope renders its code + message (e.g. 503 blocked, 400 invalid_url), 422 includes error.details, and Retry-After (including HTTP-date or garbage values) is parsed to a bounded retry delay instead of setTimeout(NaN). Specific 401/402/429 mappings still win. (#2478)

Changed

  • The bundled offline docs (get_doc/list_docs/search_docs) are regenerated: 96 pages, now including the TypeScript SDK usage and webhooks reference pages and dropping removed content. (#2469)

[0.1.3] - 2026-07-16

Changed (BREAKING)

  • The MCP tool web-search is renamed to web_search (snake_case, consistent with every other tool). MCP clients bound to the old web-search name must update. (#2421)

Changed

  • The 401 error message now references BYTEKIT_API_KEY; --base-url now falls back to the BYTEKIT_BASE_URL environment variable when no flag is passed. (#2421)

Fixed

  • The bundled offline docs (get_doc/list_docs/search_docs) now teach the correct @hunt-labs/bytekit-sdk install/import name instead of the dead @bytekit/sdk scope. (#2449)

[0.1.2] - 2026-07-15

Fixed

  • scrape_url now reads the snake_case envelope fields the API actually returns (id, final_url, status_code); the published 0.1.1 tarball read the old camelCase fields (scrapeId, finalUrl, statusCode), so every metadata line rendered undefined (#2419).
  • The MCP handshake serverInfo.version is now derived from package.json instead of a hardcoded literal, so it can never drift from the published version (#2419).

Changed

  • Compiled test code and fixtures are excluded from the published tarball — only runtime dist/** (plus README.md, LICENSE, and now CHANGELOG.md) ships (#2419).

[0.1.1] - 2026-07 [DEPRECATED]

Note

  • Superseded by 0.1.2. This release shipped stale scrape_url metadata rendering (camelCase fields the API no longer returned) and included compiled test code in the tarball. Upgrade to 0.1.2 or later.

[0.1.0] - 2026-07

Added

  • Initial release of the ByteKit MCP server, exposing ByteKit web-data tools (scrape, screenshot, search, docs) to AI agents over the Model Context Protocol, as a local stdio server or hosted.

On this page

@hunt-labs/bytekit-sdk[Unreleased][0.11.9] - 2026-10-02Changed[0.11.8] - 2026-09-30Changed[0.11.7] - 2026-09-29Changed[0.11.6] - 2026-09-24Changed[0.11.5] - 2026-09-22Changed[0.11.4] - 2026-09-21Changed[0.11.3] - 2026-09-21Changed[0.11.2] - 2026-09-18Changed[0.11.1] - 2026-09-17Changed[0.11.0] - 2026-09-14AddedChanged[0.10.0] - 2026-09-11Changed[0.9.4] - 2026-09-06 [NEVER PUBLISHED]AddedFixedChanged[0.9.3] - 2026-09-04Changed[0.9.2] - 2026-09-03Changed[0.9.1] - 2026-08-25FixedChanged[0.9.0] - 2026-08-20RemovedChanged[0.8.0] - 2026-08-19RemovedChangedFixed[0.7.4] - 2026-08-19Changed[0.7.3] - 2026-08-18Changed[0.7.2] - 2026-08-17Changed[0.7.1] - 2026-08-11ChangedAddedFixedChanged[0.7.0] - 2026-08-11AddedUnchanged[0.6.2] - 2026-08-11 [NEVER PUBLISHED]DocumentationChanged[0.6.1] - 2026-08-10Changed[0.6.0] - 2026-08-07AddedFixedChanged[0.5.0] - 2026-07-29AddedChanged[0.4.0] - 2026-07-24AddedChanged (BREAKING at compile time)FixedFixed (BREAKING at compile time)[0.3.0] - 2026-07-21Removed (BREAKING)FixedAddedChanged (BREAKING)[0.2.2] - 2026-07-16Changed[0.2.1] - 2026-07-16 [NEVER PUBLISHED]ChangedFixedRemoved[0.2.0] - 2026-07-15Changed (breaking)ChangedAdded[0.1.0] - 2026-04-24Addedbytekit-sdkUnreleased0.9.2Changed0.9.1Changed0.9.0Changed0.8.2Changed0.8.1Fixed0.8.0AddedChanged0.7.0AddedChanged0.6.0Changed0.5.4 [NEVER PUBLISHED]AddedFixedChanged0.5.3Changed0.5.2ChangedFixedChanged0.5.1FixedChanged0.5.0RemovedChanged0.4.0RemovedChangedFixed0.3.9ChangedFixedChanged0.3.8AddedFixed0.3.7 [NEVER PUBLISHED]Documentation0.3.6DocumentationChanged0.3.5BreakingFixedChangedDocumentation0.3.4Breaking0.3.3 [NEVER PUBLISHED]BreakingSecurityFixedChangedInternal0.3.2AddedFixed0.3.1 [NEVER PUBLISHED]FixedDocumentation0.3.0BreakingAdded@hunt-labs/bytekit-cli[Unreleased][0.10.1] - 2026-09-24Changed[0.10.0] - 2026-09-14Added[0.9.4] - 2026-09-10Added[0.9.3] - 2026-09-03ChangedFixed[0.9.2] - 2026-08-31Changed[0.9.1] - 2026-08-27Fixed[0.9.0] - 2026-08-20Removed[0.8.1] - 2026-08-13FixedAddedChanged[0.8.0] - 2026-08-11Added[0.7.1] - 2026-08-10Changed[0.7.0] - 2026-08-07BreakingFixed[0.6.0] - 2026-07-29BreakingAdded[0.5.1] - 2026-07-28 [NEVER PUBLISHED]Fixed[0.5.0] - 2026-07-24NoteAddedFixedDocumented[0.4.0] - 2026-07-23 [NEVER PUBLISHED]Removed (BREAKING)[0.3.1] - 2026-07-22 [NEVER PUBLISHED]FixedAdded[0.3.0] - 2026-07-21Removed (BREAKING)FixedAddedChanged (BREAKING)[0.2.2] - 2026-07-16Changed[0.2.1] - 2026-07-16 [NEVER PUBLISHED]ChangedRemoved[0.2.0] - 2026-07-15AddedFixedChanged[0.1.0] - 2026-07-13Added@hunt-labs/bytekit-mcp[Unreleased][0.3.17] - 2026-10-02Changed[0.3.16] - 2026-10-01Changed[0.3.15] - 2026-09-30FixedChanged[0.3.14] - 2026-09-29Changed[0.3.13] - 2026-09-24Changed[0.3.12] - 2026-09-23Changed[0.3.11] - 2026-09-21Changed[0.3.10] - 2026-09-21Changed[0.3.9] - 2026-09-18Changed[0.3.8] - 2026-09-17Changed[0.3.7] - 2026-09-14AddedChanged[0.3.6] - 2026-09-12Changed[0.3.5] - 2026-09-06Changed[0.3.4] - 2026-09-04Changed[0.3.3] - 2026-09-03ChangedFixedChanged[0.3.2] - 2026-08-31Changed[0.3.1] - 2026-08-25Changed[0.3.0] - 2026-08-20RemovedChanged[0.2.8] - 2026-08-19Changed[0.2.7] - 2026-08-18Changed[0.2.6] - 2026-08-12ChangedAddedFixed[0.2.5] - 2026-08-11Added[0.2.4] - 2026-08-10Changed[0.2.3] - 2026-08-07AddedFixed[0.2.2] - 2026-07-29FixedChanged[0.2.1] - 2026-07-24FixedAdded[0.2.0] - 2026-07-21Removed (BREAKING)FixedChanged[0.1.3] - 2026-07-16Changed (BREAKING)ChangedFixed[0.1.2] - 2026-07-15FixedChanged[0.1.1] - 2026-07 [DEPRECATED]Note[0.1.0] - 2026-07Added