Latest Results
Cache the istr for header names missing from hdrs in the C parser (#13887)
<!-- Thank you for your contribution! -->
The C parser returns the shared `hdrs` istr for a known header name, but
decodes every other name into a fresh `str`, which `CIMultiDict` then
turns into an `istr` again whenever the key is read. This keeps a
module-level cache from the raw name bytes to the `istr` built for it,
so a name missing from `hdrs` (vendor headers such as `CF-Ray` or
`X-Amzn-Trace-Id`, custom application headers) costs one dict lookup
after the first time it is seen.
The cache is bounded so a peer cannot grow it: at most 512 names, each
at most 64 bytes, and it is emptied when full rather than refusing new
entries, so junk names cannot keep it occupied. Longer names skip it and
are decoded as before.
Measured with callgrind on 3.14 (multidict master), instructions per
parsed request against master (which includes the extended `hdrs` list
from #13886):
| request | parse | parse + `items()` | parse + 3 `get()` |
| --- | ---: | ---: | ---: |
| behind Cloudflare/AWS, 5 vendor names | 68,232 → 64,837 (-5.0%) |
103,835 → 97,356 (-6.2%) | 81,824 → 78,371 (-4.2%) |
| Chrome page load, all names now in `hdrs` | 83,942 → 83,815 | 132,376
→ 132,274 | 94,424 → 94,260 |
| behind nginx, all names now in `hdrs` | 55,207 → 55,141 | 82,332 →
82,296 | 65,709 → 65,598 |
| curl, all names known (control) | 40,758 → 40,741 | 56,752 → 56,765 |
54,358 → 54,275 |
Before #13886 landed, the Chrome request (10 of its 15 names unknown
then) was 5.0% cheaper to parse and 7.0% cheaper with `items()`; the
cache matters for whatever `hdrs` does not list.
Returning an `istr` without the cache was measured too and is worse than
master (browser parse +12%), so the saving comes from reuse, not from
the type.
No. Header names missing from `hdrs` reach `CIMultiDict` as an `istr`
instead of a `str`, which it converts to anyway, and keep their
spelling.
No, it is a dict lookup in `find_header()` plus two constants.
Related to #13886 (merged), which added more names to `hdrs`.
- [x] I think the code is well written
- [x] Unit tests for the changes exist
- [ ] Documentation reflects the changes: N/A, internal to the C parser
- [x] If you provide code modification, please add yourself to
`CONTRIBUTORS.txt`: already listed
- [x] Add a new news fragment into the `CHANGES/` folder
Drafted with Claude Opus 5.5 (Claude Code); reviewed by @asvetlov.
<details>
<summary>Agent run details (optional, for reviewers)</summary>
- Built with Cython against multidict master (7.0.1.dev0).
- 3.14.7: `pytest --numprocesses=8`: 5629 passed, 94 skipped, 17 xfailed
(after syncing with master).
- 3.14.7, `AIOHTTP_NO_EXTENSIONS=1 pytest tests/test_http_parser.py
tests/test_web_functional.py tests/test_client_functional.py`: 947
passed, 51 skipped.
- 3.14.7t (GIL disabled): `tests/test_http_parser.py` passes except one
brotli test (brotli uninstalled locally because its 3.14t wheel fails to
import); a stress run of 8 threads each parsing 20,000 responses with
per-thread, shared and differently cased unknown names, forcing repeated
cache clears, returned every name with the right spelling.
- New tests: the same unknown name is returned as the same object; a
65-byte name is not cached; the cache is emptied after enough distinct
names.
- pre-commit: all hooks pass except flake8, whose hook environment fails
to load `flake8-requirements` on 3.14 (`pkg_resources`); flake8 run
directly is clean. mypy shows no new errors in the changed files.
- Variants measured before choosing this one: `istr` without a cache
(worse than master, above) and this cache.
- Measurement: callgrind, `PYTHONHASHSEED=0`, instrumentation limited to
a loop of `HttpRequestParserC.feed_data()` on one request, (2000 runs -
1000 runs) / 1000.
</details>
(cherry picked from commit abde3a8ef957afa1f0a2cd30ea0bf0b80050d5b6)asvetlov:patchback/backports/3.15/abde3a8ef957afa1f0a2cd30ea0bf0b80050d5b6/pr-13887 Cache the istr for header names missing from hdrs in the C parser (#13887)
<!-- Thank you for your contribution! -->
The C parser returns the shared `hdrs` istr for a known header name, but
decodes every other name into a fresh `str`, which `CIMultiDict` then
turns into an `istr` again whenever the key is read. This keeps a
module-level cache from the raw name bytes to the `istr` built for it,
so a name missing from `hdrs` (vendor headers such as `CF-Ray` or
`X-Amzn-Trace-Id`, custom application headers) costs one dict lookup
after the first time it is seen.
The cache is bounded so a peer cannot grow it: at most 512 names, each
at most 64 bytes, and it is emptied when full rather than refusing new
entries, so junk names cannot keep it occupied. Longer names skip it and
are decoded as before.
Measured with callgrind on 3.14 (multidict master), instructions per
parsed request against master (which includes the extended `hdrs` list
from #13886):
| request | parse | parse + `items()` | parse + 3 `get()` |
| --- | ---: | ---: | ---: |
| behind Cloudflare/AWS, 5 vendor names | 68,232 → 64,837 (-5.0%) |
103,835 → 97,356 (-6.2%) | 81,824 → 78,371 (-4.2%) |
| Chrome page load, all names now in `hdrs` | 83,942 → 83,815 | 132,376
→ 132,274 | 94,424 → 94,260 |
| behind nginx, all names now in `hdrs` | 55,207 → 55,141 | 82,332 →
82,296 | 65,709 → 65,598 |
| curl, all names known (control) | 40,758 → 40,741 | 56,752 → 56,765 |
54,358 → 54,275 |
Before #13886 landed, the Chrome request (10 of its 15 names unknown
then) was 5.0% cheaper to parse and 7.0% cheaper with `items()`; the
cache matters for whatever `hdrs` does not list.
Returning an `istr` without the cache was measured too and is worse than
master (browser parse +12%), so the saving comes from reuse, not from
the type.
No. Header names missing from `hdrs` reach `CIMultiDict` as an `istr`
instead of a `str`, which it converts to anyway, and keep their
spelling.
No, it is a dict lookup in `find_header()` plus two constants.
Related to #13886 (merged), which added more names to `hdrs`.
- [x] I think the code is well written
- [x] Unit tests for the changes exist
- [ ] Documentation reflects the changes: N/A, internal to the C parser
- [x] If you provide code modification, please add yourself to
`CONTRIBUTORS.txt`: already listed
- [x] Add a new news fragment into the `CHANGES/` folder
Drafted with Claude Opus 5.5 (Claude Code); reviewed by @asvetlov.
<details>
<summary>Agent run details (optional, for reviewers)</summary>
- Built with Cython against multidict master (7.0.1.dev0).
- 3.14.7: `pytest --numprocesses=8`: 5629 passed, 94 skipped, 17 xfailed
(after syncing with master).
- 3.14.7, `AIOHTTP_NO_EXTENSIONS=1 pytest tests/test_http_parser.py
tests/test_web_functional.py tests/test_client_functional.py`: 947
passed, 51 skipped.
- 3.14.7t (GIL disabled): `tests/test_http_parser.py` passes except one
brotli test (brotli uninstalled locally because its 3.14t wheel fails to
import); a stress run of 8 threads each parsing 20,000 responses with
per-thread, shared and differently cased unknown names, forcing repeated
cache clears, returned every name with the right spelling.
- New tests: the same unknown name is returned as the same object; a
65-byte name is not cached; the cache is emptied after enough distinct
names.
- pre-commit: all hooks pass except flake8, whose hook environment fails
to load `flake8-requirements` on 3.14 (`pkg_resources`); flake8 run
directly is clean. mypy shows no new errors in the changed files.
- Variants measured before choosing this one: `istr` without a cache
(worse than master, above) and this cache.
- Measurement: callgrind, `PYTHONHASHSEED=0`, instrumentation limited to
a loop of `HttpRequestParserC.feed_data()` on one request, (2000 runs -
1000 runs) / 1000.
</details>
(cherry picked from commit abde3a8ef957afa1f0a2cd30ea0bf0b80050d5b6)asvetlov:patchback/backports/3.14/abde3a8ef957afa1f0a2cd30ea0bf0b80050d5b6/pr-13887 Cache the istr for header names missing from hdrs in the C parser (#13887)
<!-- Thank you for your contribution! -->
## What do these changes do?
The C parser returns the shared `hdrs` istr for a known header name, but
decodes every other name into a fresh `str`, which `CIMultiDict` then
turns into an `istr` again whenever the key is read. This keeps a
module-level cache from the raw name bytes to the `istr` built for it,
so a name missing from `hdrs` (vendor headers such as `CF-Ray` or
`X-Amzn-Trace-Id`, custom application headers) costs one dict lookup
after the first time it is seen.
The cache is bounded so a peer cannot grow it: at most 512 names, each
at most 64 bytes, and it is emptied when full rather than refusing new
entries, so junk names cannot keep it occupied. Longer names skip it and
are decoded as before.
Measured with callgrind on 3.14 (multidict master), instructions per
parsed request against master (which includes the extended `hdrs` list
from #13886):
| request | parse | parse + `items()` | parse + 3 `get()` |
| --- | ---: | ---: | ---: |
| behind Cloudflare/AWS, 5 vendor names | 68,232 → 64,837 (-5.0%) |
103,835 → 97,356 (-6.2%) | 81,824 → 78,371 (-4.2%) |
| Chrome page load, all names now in `hdrs` | 83,942 → 83,815 | 132,376
→ 132,274 | 94,424 → 94,260 |
| behind nginx, all names now in `hdrs` | 55,207 → 55,141 | 82,332 →
82,296 | 65,709 → 65,598 |
| curl, all names known (control) | 40,758 → 40,741 | 56,752 → 56,765 |
54,358 → 54,275 |
Before #13886 landed, the Chrome request (10 of its 15 names unknown
then) was 5.0% cheaper to parse and 7.0% cheaper with `items()`; the
cache matters for whatever `hdrs` does not list.
Returning an `istr` without the cache was measured too and is worse than
master (browser parse +12%), so the saving comes from reuse, not from
the type.
## Are there changes in behavior for the user?
No. Header names missing from `hdrs` reach `CIMultiDict` as an `istr`
instead of a `str`, which it converts to anyway, and keep their
spelling.
## Is it a substantial burden for the maintainers to support this?
No, it is a dict lookup in `find_header()` plus two constants.
## Related issue number
Related to #13886 (merged), which added more names to `hdrs`.
## Checklist
- [x] I think the code is well written
- [x] Unit tests for the changes exist
- [ ] Documentation reflects the changes: N/A, internal to the C parser
- [x] If you provide code modification, please add yourself to
`CONTRIBUTORS.txt`: already listed
- [x] Add a new news fragment into the `CHANGES/` folder
Drafted with Claude Opus 5.5 (Claude Code); reviewed by @asvetlov.
<details>
<summary>Agent run details (optional, for reviewers)</summary>
- Built with Cython against multidict master (7.0.1.dev0).
- 3.14.7: `pytest --numprocesses=8`: 5629 passed, 94 skipped, 17 xfailed
(after syncing with master).
- 3.14.7, `AIOHTTP_NO_EXTENSIONS=1 pytest tests/test_http_parser.py
tests/test_web_functional.py tests/test_client_functional.py`: 947
passed, 51 skipped.
- 3.14.7t (GIL disabled): `tests/test_http_parser.py` passes except one
brotli test (brotli uninstalled locally because its 3.14t wheel fails to
import); a stress run of 8 threads each parsing 20,000 responses with
per-thread, shared and differently cased unknown names, forcing repeated
cache clears, returned every name with the right spelling.
- New tests: the same unknown name is returned as the same object; a
65-byte name is not cached; the cache is emptied after enough distinct
names.
- pre-commit: all hooks pass except flake8, whose hook environment fails
to load `flake8-requirements` on 3.14 (`pkg_resources`); flake8 run
directly is clean. mypy shows no new errors in the changed files.
- Variants measured before choosing this one: `istr` without a cache
(worse than master, above) and this cache.
- Measurement: callgrind, `PYTHONHASHSEED=0`, instrumentation limited to
a loop of `HttpRequestParserC.feed_data()` on one request, (2000 runs -
1000 runs) / 1000.
</details> Latest Branches
0%
muhammad-a-dev:fix-content-encoding-case-13894 +16%
asvetlov:patchback/backports/3.15/abde3a8ef957afa1f0a2cd30ea0bf0b80050d5b6/pr-13887 0%
asvetlov:patchback/backports/3.14/abde3a8ef957afa1f0a2cd30ea0bf0b80050d5b6/pr-13887 © 2026 CodSpeed Technology