Triage and debug go.dev/issue/78438
#78438 (HTTP gzip flake)
/home/claude/golang_go_78438/go · Wed 26 Aug 2026 15:07:52 · 20 messages399e501a-6ca0-43ff-9dd2-f1fc86d93995
You Tue 25 Aug 2026 08:12:24
Assistant Tue 25 Aug 2026 08:12:26
I'll start by fetching the issue.
Ran tools (1 ToolSearch, 16 Bash, 2 WebFetch)
ToolSearch select:WebFetch,WebSearch
select:WebFetch,WebSearch
5
Bash ls /home/claude/golang_go_78438 2>/dev/null | head -50; echo "---"; ls -a /home/claude/golang_go_78438 2>/dev/null | hea...
ls /home/claude/golang_go_78438 2>/dev/null | head -50; echo "---"; ls -a /home/claude/golang_go_78438 2>/dev/null | head -20
List working directory contents
env.sh go --- . .. env.sh go
WebFetch https://go.dev/issue/78438
https://go.dev/issue/78438
Reproduce the full issue text verbatim: title, author, labels, body (including any code and error output), and all comments.
REDIRECT DETECTED: The URL redirects to a location that was not fetched automatically. Original URL: https://go.dev/issue/78438 Redirect URL (from the server's Location header — server-supplied, not verified): https://github.com/golang/go/issues/78438 Status: 302 Found To complete your request, I need to fetch content from the redirected URL. Please use WebFetch again with these parameters: - url: "https://github.com/golang/go/issues/78438" - prompt: "Reproduce the full issue text verbatim: title, author, labels, body (including any code and error output), and all comments."
Bash cat /home/claude/golang_go_78438/env.sh; echo "---"; ls /home/claude/golang_go_78438/go | head; echo "---"; cd /home/cla...
cat /home/claude/golang_go_78438/env.sh; echo "---"; ls /home/claude/golang_go_78438/go | head; echo "---"; cd /home/claude/golang_go_78438/go && git log --oneline -3 2>&1 | head
Inspect env and go repo
export WORK=/home/claude/golang_go_78438 export GOROOT=$WORK/go export GOPATH=$WORK export GOBIN=$WORK/bin export PATH=$GOBIN:/home/claude/.local/bin:/home/claude/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/usr/games:/usr/local/games:/snap/bin --- api bin codereview.cfg CONTRIBUTING.md doc go.env lib LICENSE misc PATENTS --- 13b477edc7 runtime: fix flaky refill detection in TestScanAllocIssue77573 72755de33c os, internal/poll: use splice for copies involving pipes 3e6ad2cd0a cmd/compile: always do rewriteCondSelectIntoMath for powers of two
WebFetch https://github.com/golang/go/issues/78438
https://github.com/golang/go/issues/78438
Reproduce the full issue text verbatim: title, author, labels, body (including any code and error output), and all comments.
# Issue Summary **Title:** net/http: TestTransportGzip/h1 failures #78438 **Author:** gopherbot **Labels:** - Automation - NeedsInvestigation **Body:** The issue was created automatically by watchflakes to track test failures. The command used was: ``` #!watchflakes default <- pkg == "net/http" && test == "TestTransportGzip/h1" ``` An example failure log is referenced from a Chromium CI build showing: ``` === RUN TestTransportGzip/h1 === PAUSE TestTransportGzip/h1 === CONT TestTransportGzip/h1 ``` The issue indicates that "TestTransportGzip/h1" test in Go's net/http package is experiencing intermittent failures that require investigation to determine if this represents a genuine bug or a duplicate of an existing issue. **Status:** Open **Project:** Test Flakes
Bash (gh issue view 78438 --repo golang/go --comments 2>&1 | head -200) || curl -s https://api.github.com/repos/golang/go/iss...
(gh issue view 78438 --repo golang/go --comments 2>&1 | head -200) || curl -s https://api.github.com/repos/golang/go/issues/78438 | head -100
Fetch issue via gh or API
/bin/bash: line 1: gh: command not found
Bash cd /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad 2>/dev/null || cd /tmp;...
cd /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad 2>/dev/null || cd /tmp; curl -sL https://api.github.com/repos/golang/go/issues/78438 -o iss.json; python3 -c "
import json
d=json.load(open('iss.json'))
print(d['title']); print(d['state']); print([l['name'] for l in d['labels']])
print(d['body'])
"Fetch issue JSON body
net/http: TestTransportGzip/h1 failures
open
['NeedsInvestigation', 'Automation']
```
#!watchflakes
default <- pkg == "net/http" && test == "TestTransportGzip/h1"
```
Issue created automatically to collect these failures.
Example ([log](https://ci.chromium.org/b/8686112892523717025)):
=== RUN TestTransportGzip/h1
=== PAUSE TestTransportGzip/h1
=== CONT TestTransportGzip/h1
— [watchflakes](https://go.dev/wiki/Watchflakes)
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; curl -sL "https://api.github.com/repos/golang/go/issues/78438/comments" -o cm.json; python3 -c " ...
cd /tmp; curl -sL "https://api.github.com/repos/golang/go/issues/78438/comments" -o cm.json; python3 -c "
import json
d=json.load(open('cm.json'))
print(len(d))
for c in d:
print('=== ', c['user']['login'], c['created_at'])
print(c['body'][:6000])
"Fetch issue comments
7
=== gopherbot 2026-03-28T21:24:54Z
Found new dashboard test flakes for:
#!watchflakes
default <- pkg == "net/http" && test == "TestTransportGzip/h1"
<details><summary>2026-03-28 04:32 gotip-solaris-amd64 go@1fd68799 net/http.TestTransportGzip/h1 [SKIP] (<a href="https://ci.chromium.org/b/8686112892523717025">log</a>)</summary>
=== RUN TestTransportGzip/h1
=== PAUSE TestTransportGzip/h1
=== CONT TestTransportGzip/h1
</details>
— [watchflakes](https://go.dev/wiki/Watchflakes)
=== JunyangShao 2026-03-31T17:32:48Z
Looks like it just timeout-ed?
@neild @nicholashusin
=== gopherbot 2026-05-19T19:55:59Z
Found new dashboard test flakes for:
#!watchflakes
default <- pkg == "net/http" && test == "TestTransportGzip/h1"
<details><summary>2026-05-18 18:14 gotip-linux-ppc64le_power8 go@2e67b189 net/http.TestTransportGzip/h1 [SKIP] (<a href="https://ci.chromium.org/b/8681423021638815105">log</a>)</summary>
=== RUN TestTransportGzip/h1
=== PAUSE TestTransportGzip/h1
=== CONT TestTransportGzip/h1
</details>
— [watchflakes](https://go.dev/wiki/Watchflakes)
=== gopherbot 2026-05-23T11:01:49Z
Found new dashboard test flakes for:
#!watchflakes
default <- pkg == "net/http" && test == "TestTransportGzip/h1"
<details><summary>2026-05-19 15:57 gotip-solaris-amd64 go@05f75fb9 net/http.TestTransportGzip/h1 [SKIP] (<a href="https://ci.chromium.org/b/8681022390813019185">log</a>)</summary>
=== RUN TestTransportGzip/h1
=== PAUSE TestTransportGzip/h1
=== CONT TestTransportGzip/h1
</details>
— [watchflakes](https://go.dev/wiki/Watchflakes)
=== gopherbot 2026-07-16T22:07:55Z
Found new dashboard test flakes for:
#!watchflakes
default <- pkg == "net/http" && test == "TestTransportGzip/h1"
<details><summary>2026-07-13 19:15 go1.27-solaris-amd64 release-branch.go1.27@96307fb5 net/http.TestTransportGzip/h1 [SKIP] (<a href="https://ci.chromium.org/b/8676093236579846177">log</a>)</summary>
=== RUN TestTransportGzip/h1
=== PAUSE TestTransportGzip/h1
=== CONT TestTransportGzip/h1
</details>
— [watchflakes](https://go.dev/wiki/Watchflakes)
=== gopherbot 2026-07-19T11:04:33Z
Found new dashboard test flakes for:
#!watchflakes
default <- pkg == "net/http" && test == "TestTransportGzip/h1"
<details><summary>2026-07-15 15:30 gotip-solaris-amd64 go@c29765a8 net/http.TestTransportGzip/h1 [SKIP] (<a href="https://ci.chromium.org/b/8675856801014264801">log</a>)</summary>
=== RUN TestTransportGzip/h1
=== PAUSE TestTransportGzip/h1
=== CONT TestTransportGzip/h1
</details>
— [watchflakes](https://go.dev/wiki/Watchflakes)
=== gopherbot 2026-08-22T03:25:24Z
Found new dashboard test flakes for:
#!watchflakes
default <- pkg == "net/http" && test == "TestTransportGzip/h1"
<details><summary>2026-08-21 21:04 gotip-solaris-amd64 go@d3edcca1 net/http.TestTransportGzip/h1 [SKIP] (<a href="https://ci.chromium.org/b/8672808403636415393">log</a>)</summary>
=== RUN TestTransportGzip/h1
=== PAUSE TestTransportGzip/h1
=== CONT TestTransportGzip/h1
</details>
— [watchflakes](https://go.dev/wiki/Watchflakes)
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /home/claude/golang_go_78438/go/src/net/http && grep -rn "TestTransportGzip" --include=*.go . | head
cd /home/claude/golang_go_78438/go/src/net/http && grep -rn "TestTransportGzip" --include=*.go . | head
Locate TestTransportGzip
transport_test.go:1203:func TestTransportGzip(t *testing.T) { run(t, testTransportGzip, http3SkippedMode) }
transport_test.go:2003:// TestTransportGzipRecursive sends a gzip quine and checks that the
transport_test.go:2007:func TestTransportGzipRecursive(t *testing.T) { run(t, testTransportGzipRecursive, http3SkippedMode) }
transport_test.go:2034:func TestTransportGzipShort(t *testing.T) { run(t, testTransportGzipShort, http3SkippedMode) }
transport_test.go:2056:func TestTransportGzipConcurrentCloseAndRead(t *testing.T) {Bash cd /tmp; curl -s -X POST "https://cr-buildbucket.appspot.com/prpc/buildbucket.v2.Builds/GetBuild" -H "Content-Type: appl...
cd /tmp; curl -s -X POST "https://cr-buildbucket.appspot.com/prpc/buildbucket.v2.Builds/GetBuild" -H "Content-Type: application/json" -H "Accept: application/json" -d '{"id":"8672808403636415393","mask":{"fields":"id,builder,status,summaryMarkdown,steps"}}' -o b.json; head -c 300 b.jsonTry buildbucket API for build info
)]}'
{"id":"8672808403636415393","builder":{"project":"golang","bucket":"ci","builder":"gotip-solaris-amd64"},"status":"FAILURE","summaryMarkdown":"* go on master ([commit d3edcca](https://go.googlesource.com/go/+/d3edcca1d26389184ceba1f6f8d224a5c31719d5))\n\nTo reproduce, try `gomote repro 86728084
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; python3 -c " ...
cd /tmp; python3 -c "
import json
s=open('b.json').read()
s=s[s.index('{'):]
d=json.loads(s)
print(d['summaryMarkdown'][:3000])
print('---STEPS---')
for st in d.get('steps',[]):
if st.get('status')!='SUCCESS':
print(st['name'], st['status'])
for l in st.get('logs',[]):
print(' ', l['name'], l.get('viewUrl'))
"Parse build info
* go on master ([commit d3edcca](https://go.googlesource.com/go/+/d3edcca1d26389184ceba1f6f8d224a5c31719d5))
To reproduce, try `gomote repro 8672808403636415393`.
Additional links for debugging:
* [go tool dist test -json output](https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672808403636415393/+/u/step/11/log/2)
---STEPS---
run tests FAILURE
run tests|go tool dist test -json FAILURE
command https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672808403636415393/+/u/step/11/log/1
output https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672808403636415393/+/u/step/11/log/2
file: dist.testjson https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672808403636415393/+/u/step/11/log/3
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672808403636415393/+/u/step/11/log/...
cd /tmp; curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672808403636415393/+/u/step/11/log/2?format=raw" -o log2.txt; ls -l log2.txt; head -c 500 log2.txt
Fetch dist test output log
-rw-rw-r-- 1 claude claude 13780 Aug 25 08:13 log2.txt Use of LUCI is subject to the Google [Terms of Service](https://policies.google.com/terms) and [Privacy Policy](https://policies.google.com/privacy) [I2026-08-22T04:02:26.070759+02:00 27821 0 sink.go:371] SinkServer: warm-up started [I2026-08-22T04:02:26.071427+02:00 27821 0 sink.go:454] SinkServer: starting HTTP server... [I2026-08-22T04:02:26.072647+02:00 27821 0 sink.go:376] SinkServer: warm-up ended [I2026-08-22T04:02:26.072927+02:00 27821 0 cmd_stream.go:577] rdb-stream: starting the test Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; grep -n "FAIL\|panic\|timed out\|net/http" log2.txt | head -60
cd /tmp; grep -n "FAIL\|panic\|timed out\|net/http" log2.txt | head -60
Search log for failures
208:FAIL net/http 600.080s 209:ok net/http/cgi 0.434s 210:ok net/http/cookiejar 0.038s 211:ok net/http/fcgi 0.238s 212:ok net/http/httptest 0.150s 213:ok net/http/httptrace 0.050s 214:ok net/http/httputil 0.729s 215:ok net/http/internal 0.217s 216:ok net/http/internal/ascii 0.030s 217:ok net/http/internal/http2 7.309s 218:ok net/http/internal/httpcommon 0.042s 219:ok net/http/pprof 5.123s Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672808403636415393/+/u/step/11/log/...
cd /tmp; curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672808403636415393/+/u/step/11/log/3?format=raw" -o log3.txt; ls -l log3.txt
Fetch dist.testjson
-rw-rw-r-- 1 claude claude 117334637 Aug 25 08:13 log3.txt Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; python3 - <<'EOF' ...
cd /tmp; python3 - <<'EOF'
import json
out=[]
with open('log3.txt') as f:
for line in f:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
s=''.join(out)
open('nethttp.txt','w').write(s)
print(len(s))
EOF
grep -n "panic: test timed out" nethttp.txt | headExtract net/http output
277513 5292:panic: test timed out after 10m0s Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; sed -n '5280,5420p' nethttp.txt
cd /tmp; sed -n '5280,5420p' nethttp.txt
Show timeout panic head
--- PASS: TestClientRedirects/h1 (0.03s)
--- PASS: TestClientRedirects/h2 (0.05s)
--- PASS: TestClientRedirects (0.00s)
--- PASS: TestRequestLimit/h2 (1.47s)
--- PASS: TestRequestLimit (0.00s)
serve_test.go:6260: http://127.0.0.1:46476: "127.0.0.1:49712"
serve_test.go:6260: https://127.0.0.1:54039: "127.0.0.1:62940"
--- PASS: TestServerShutdown/h1 (1.04s)
--- PASS: TestServerShutdown/h2 (1.06s)
--- PASS: TestServerShutdown (0.00s)
2026/08/22 04:04:32 httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0x1219fed3e300 127.0.0.1:47089 in state active
panic: test timed out after 10m0s
running tests:
TestTransportGzip/h1 (9m58s)
goroutine 19549 gp=0x1219ff816000 m=33 mp=0x1219feb4d008 [running]:
panic({0xee8e38?, 0x121a00654030?})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/panic.go:878 +0x151 fp=0x1219fea79f10 sp=0x1219fea79e68 pc=0x486f31
testing.(*M).startAlarm.func1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2966 +0x34a fp=0x1219fea79fe0 sp=0x1219fea79f10 pc=0x523fea
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fea79fe8 sp=0x1219fea79fe0 pc=0x48f641
created by time.goFunc
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/time/sleep.go:187 +0x2d
goroutine 1 gp=0x1219fe8ac1e0 m=nil [chan receive, 9 minutes]:
runtime.gopark(0x1219fe93aee0?, 0x1219fe9218e8?, 0x6c?, 0x8b?, 0xfecd60?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219fe921858 sp=0x1219fe921838 pc=0x487406
runtime.chanrecv(0x1219fef00d00, 0x0, 0x1)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/chan.go:659 +0x4bd fp=0x1219fe9218d0 sp=0x1219fe921858 pc=0x41763d
runtime.chanrecv1(0x1219fea401c0?, 0xa05460?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/chan.go:501 +0x12 fp=0x1219fe9218f8 sp=0x1219fe9218d0 pc=0x417152
testing.tRunner.func1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2149 +0x425 fp=0x1219fe921a90 sp=0x1219fe9218f8 pc=0x523305
testing.tRunner(0x1219fe9c0248, 0x1219fe921b98)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2206 +0x123 fp=0x1219fe921ae0 sp=0x1219fe921a90 pc=0x51dca3
testing.runTests({0x0, 0x0}, {0x9a3061, 0x8}, 0x1219fe9a21c8, {0xfe79a0, 0x26e, 0x26e}, {0xc29a5f605e9ae13c, 0x8bb2f65872, ...})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2747 +0x4d0 fp=0x1219fe921bc8 sp=0x1219fe921ae0 pc=0x520150
testing.(*M).Run(0x1219fe9ba3c0)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2607 +0x68f fp=0x1219fe921e08 sp=0x1219fe921bc8 pc=0x51edaf
net/http_test.TestMain(0x100ed40?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/main_test.go:24 +0x27 fp=0x1219fe921e28 sp=0x1219fe921e08 pc=0x857ce7
main.main()
_testmain.go:1360 +0xa5 fp=0x1219fe921eb8 sp=0x1219fe921e28 pc=0x92d985
runtime.main()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:302 +0x41b fp=0x1219fe921fe0 sp=0x1219fe921eb8 pc=0x44febb
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fe921fe8 sp=0x1219fe921fe0 pc=0x48f641
goroutine 2 gp=0x1219fe8acd20 m=nil [force gc (idle), 3 minutes]:
runtime.gopark(0x0?, 0x0?, 0x0?, 0x0?, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219fe90efa8 sp=0x1219fe90ef88 pc=0x487406
runtime.goparkunlock(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:480
runtime.forcegchelper()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:387 +0xa5 fp=0x1219fe90efe0 sp=0x1219fe90efa8 pc=0x450165
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fe90efe8 sp=0x1219fe90efe0 pc=0x48f641
created by runtime.init.7 in goroutine 1
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:375 +0x1a
goroutine 3 gp=0x1219fe8ad2c0 m=nil [GC sweep wait]:
runtime.gopark(0x1?, 0x0?, 0x0?, 0x0?, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219fe90f788 sp=0x1219fe90f768 pc=0x487406
runtime.goparkunlock(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:480
runtime.bgsweep(0x1219fe938000)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgcsweep.go:324 +0x151 fp=0x1219fe90f7c8 sp=0x1219fe90f788 pc=0x438e11
runtime.gcenable.gowrap1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:214 +0x17 fp=0x1219fe90f7e0 sp=0x1219fe90f7c8 pc=0x47d2d7
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fe90f7e8 sp=0x1219fe90f7e0 pc=0x48f641
created by runtime.gcenable in goroutine 1
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:214 +0x66
goroutine 4 gp=0x1219fe8ad4a0 m=nil [GC scavenge wait]:
runtime.gopark(0x10a12e?, 0x799cd?, 0x0?, 0x0?, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219fe90ff78 sp=0x1219fe90ff58 pc=0x487406
runtime.goparkunlock(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:480
runtime.(*scavengerState).park(0xfed560)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgcscavenge.go:425 +0x49 fp=0x1219fe90ffa8 sp=0x1219fe90ff78 pc=0x436a29
runtime.bgscavenge(0x1219fe938000)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgcscavenge.go:658 +0x59 fp=0x1219fe90ffc8 sp=0x1219fe90ffa8 pc=0x436f19
runtime.gcenable.gowrap2()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:215 +0x17 fp=0x1219fe90ffe0 sp=0x1219fe90ffc8 pc=0x47d297
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fe90ffe8 sp=0x1219fe90ffe0 pc=0x48f641
created by runtime.gcenable in goroutine 1
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:215 +0xa5
goroutine 5 gp=0x1219fe8ada40 m=nil [GOMAXPROCS updater (idle), 10 minutes]:
runtime.gopark(0x0?, 0x0?, 0x0?, 0x0?, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219fe90e788 sp=0x1219fe90e768 pc=0x487406
runtime.goparkunlock(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:480
runtime.updateMaxProcsGoroutine()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:7147 +0xe7 fp=0x1219fe90e7e0 sp=0x1219fe90e788 pc=0x45cf67
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fe90e7e8 sp=0x1219fe90e7e0 pc=0x48f641
created by runtime.defaultGOMAXPROCSUpdateEnable in goroutine 1
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:7135 +0x33
goroutine 6 gp=0x1219fe9501e0 m=nil [finalizer wait, 9 minutes]:
runtime.gopark(0x0?, 0xf70688?, 0x0?, 0x40?, 0x2000000020?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219fe910620 sp=0x1219fe910600 pc=0x487406
runtime.runFinalizers()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mfinal.go:210 +0x105 fp=0x1219fe9107e0 sp=0x1219fe910620 pc=0x42a3a5
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fe9107e8 sp=0x1219fe9107e0 pc=0x48f641
created by runtime.createfing in goroutine 1
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mfinal.go:172 +0x3d
goroutine 18 gp=0x1219fe9da000 m=nil [cleanup wait, 9 minutes]:
runtime.gopark(0x610914?, 0x7fffbecd8600?, 0x90?, 0xa7?, 0x41ca35?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219fe90a760 sp=0x1219fe90a740 pc=0x487406
runtime.goparkunlock(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:480
runtime.(*cleanupQueue).dequeue(0xfed900)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mcleanup.go:522 +0xd7 fp=0x1219fe90a7a0 sp=0x1219fe90a760 pc=0x427477
runtime.runCleanups()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mcleanup.go:718 +0x45 fp=0x1219fe90a7e0 sp=0x1219fe90a7a0 pc=0x427aa5
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fe90a7e8 sp=0x1219fe90a7e0 pc=0x48f641
created by runtime.(*cleanupQueue).createGs in goroutine 1
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mcleanup.go:672 +0xa5
goroutine 7 gp=0x1219fe950780 m=nil [GC worker (idle), 9 minutes]:
runtime.gopark(0x10102e0?, 0x1?, 0x73?, 0x49?, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219feef0f40 sp=0x1219feef0f20 pc=0x487406
runtime.gcBgMarkWorker(0x1219fe9440e0)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:1807 +0xe9 fp=0x1219feef0fc8 sp=0x1219feef0f40 pc=0x42d7a9
runtime.gcBgMarkStartWorkers.gowrap1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:1711 +0x17 fp=0x1219feef0fe0 sp=0x1219feef0fc8 pc=0x47d777
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219feef0fe8 sp=0x1219feef0fe0 pc=0x48f641
created by runtime.gcBgMarkStartWorkers in goroutine 22
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:1711 +0xfc
goroutine 8 gp=0x1219fe950960 m=nil [GC worker (idle), 9 minutes]:
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; grep -n "net/http_test\.\|net/http\.(\*" nethttp.txt | sed -n '1,200p'
cd /tmp; grep -n "net/http_test\.\|net/http\.(\*" nethttp.txt | sed -n '1,200p'
Find test-related goroutine frames
4833: serve_test.go:7228: log: http: superfluous response.WriteHeader call from net/http_test.testWriteHeaderSwitchingProtocols.func1 (serve_test.go:7245)
5321:net/http_test.TestMain(0x100ed40?)
5539:net/http.(*chunkWriter).Write(0x121a00f7f8a8, {0x121a027e8000, 0x1000a4, 0x120000})
5543:net/http.(*response).write(0x121a00f7f860, 0x1000a4, {0x121a027e8000, 0x1000a4, 0x120000}, {0x0, 0x0})
5545:net/http.(*response).Write(0xe6b0d0?, {0x121a027e8000?, 0x1219ff8c5920?, 0x411989?})
5553:net/http_test.testTransportGzip.func1.deferwrap1()
5555:net/http_test.testTransportGzip.func1({0xf69e00, 0x121a00f7f860}, 0x121a011617c0)
5561:net/http.(*conn).serve(0x121a012bd440, {0xf6ad78, 0x1219ff547170})
5563:net/http.(*Server).Serve.gowrap3()
5567:created by net/http.(*Server).Serve in goroutine 8900
5601:net/http_test.(*clientServerTest).close(0x1219ff495400)
5603:net/http_test.newClientServerTest.func2()
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; sed -n '5500,5640p' nethttp.txt
cd /tmp; sed -n '5500,5640p' nethttp.txt
Show blocked goroutines
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fed0ffe8 sp=0x1219fed0ffe0 pc=0x48f641
created by runtime.gcBgMarkStartWorkers in goroutine 22
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:1711 +0xfc
goroutine 903 gp=0x1219fee6a5a0 m=nil [cleanup wait, 9 minutes]:
runtime.gopark(0x610914?, 0x7fffbf141000?, 0x90?, 0xa7?, 0x41ca35?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219feb2a760 sp=0x1219feb2a740 pc=0x487406
runtime.goparkunlock(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:480
runtime.(*cleanupQueue).dequeue(0xfed900)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mcleanup.go:522 +0xd7 fp=0x1219feb2a7a0 sp=0x1219feb2a760 pc=0x427477
runtime.runCleanups()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mcleanup.go:718 +0x45 fp=0x1219feb2a7e0 sp=0x1219feb2a7a0 pc=0x427aa5
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219feb2a7e8 sp=0x1219feb2a7e0 pc=0x48f641
created by runtime.(*cleanupQueue).createGs in goroutine 902
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/mcleanup.go:672 +0xa5
goroutine 8885 gp=0x1219feb10d20 m=nil [IO wait, 9 minutes]:
runtime.gopark(0x1219ff8c53f8?, 0x4493ca?, 0x0?, 0xda?, 0x3?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219ff8c53d0 sp=0x1219ff8c53b0 pc=0x487406
runtime.netpollblock(0x7fffbf2bda00?, 0x77?, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/netpoll.go:575 +0xf7 fp=0x1219ff8c5408 sp=0x1219ff8c53d0 pc=0x448a77
internal/poll.runtime_pollWait(0x7fffbf2bda00, 0x77)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/netpoll.go:351 +0x8b fp=0x1219ff8c5428 sp=0x1219ff8c5408 pc=0x48668b
internal/poll.(*pollDesc).wait(0x1219ff5adb00?, 0x121a02884f97?, 0x0)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/internal/poll/fd_poll_runtime.go:84 +0x27 fp=0x1219ff8c5460 sp=0x1219ff8c5428 pc=0x4beda7
internal/poll.(*pollDesc).waitWrite(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/internal/poll/fd_poll_runtime.go:93
internal/poll.(*FD).Write(0x1219ff5adb00, {0x121a027e8f97, 0xff10d, 0x11f069})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/internal/poll/fd_unix.go:395 +0x30a fp=0x1219ff8c5548 sp=0x1219ff8c5460 pc=0x4c078a
net.(*netFD).Write(0x1219ff5adb00, {0x121a027e8f97?, 0x0?, 0x0?})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/fd_posix.go:109 +0x25 fp=0x1219ff8c5590 sp=0x1219ff8c5548 pc=0x626085
net.(*conn).Write(0x1219fed3e300, {0x121a027e8f97?, 0x1219ff99b000?, 0x1219ff6af980?})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/net.go:208 +0x45 fp=0x1219ff8c55d8 sp=0x1219ff8c5590 pc=0x630da5
net/http.checkConnErrorWriter.Write({0x121a02560c40?}, {0x121a027e8f97?, 0x813322e9a3b91482?, 0x98483a12874577d5?})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:4263 +0x26 fp=0x1219ff8c5628 sp=0x1219ff8c55d8 pc=0x76ca66
bufio.(*Writer).Write(0x121a02560c40, {0x121a027e8000?, 0x1000a4?, 0x120000?})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/bufio/bufio.go:682 +0xec fp=0x1219ff8c5688 sp=0x1219ff8c5628 pc=0x526f0c
net/http.(*chunkWriter).Write(0x121a00f7f8a8, {0x121a027e8000, 0x1000a4, 0x120000})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:392 +0xff fp=0x1219ff8c56f8 sp=0x1219ff8c5688 pc=0x75dc1f
bufio.(*Writer).Write(0x1219ff0596c0, {0x121a027e8000?, 0x5?, 0x0?})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/bufio/bufio.go:682 +0xec fp=0x1219ff8c5758 sp=0x1219ff8c56f8 pc=0x526f0c
net/http.(*response).write(0x121a00f7f860, 0x1000a4, {0x121a027e8000, 0x1000a4, 0x120000}, {0x0, 0x0})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:1690 +0x1be fp=0x1219ff8c5858 sp=0x1219ff8c5758 pc=0x763f1e
net/http.(*response).Write(0xe6b0d0?, {0x121a027e8000?, 0x1219ff8c5920?, 0x411989?})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:1653 +0x2a fp=0x1219ff8c58a0 sp=0x1219ff8c5858 pc=0x763cca
bytes.(*Buffer).WriteTo(0x121a009c17d0, {0xf66940?, 0x121a00f7f860?})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/bytes/buffer.go:279 +0x7b fp=0x1219ff8c58d8 sp=0x1219ff8c58a0 pc=0x50673b
io.copyBuffer({0xf66940, 0x121a00f7f860}, {0xf66260, 0x121a009c17d0}, {0x0, 0x0, 0x0})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/io/io.go:411 +0x9d fp=0x1219ff8c5950 sp=0x1219ff8c58d8 pc=0x4b611d
io.Copy(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/io/io.go:388
net/http_test.testTransportGzip.func1.deferwrap1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/transport_test.go:1223 +0x2b fp=0x1219ff8c5998 sp=0x1219ff8c5950 pc=0x8fa64b
net/http_test.testTransportGzip.func1({0xf69e00, 0x121a00f7f860}, 0x121a011617c0)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/transport_test.go:1234 +0x407 fp=0x1219ff8c5aa0 sp=0x1219ff8c5998 pc=0x8fa387
net/http.HandlerFunc.ServeHTTP(0x10?, {0xf69e00?, 0x121a00f7f860?}, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:2336 +0x29 fp=0x1219ff8c5ac8 sp=0x1219ff8c5aa0 pc=0x766b49
net/http.serverHandler.ServeHTTP({0x121a02560c00?}, {0xf69e00?, 0x121a00f7f860?}, 0x1?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:3421 +0xbc fp=0x1219ff8c5b18 sp=0x1219ff8c5ac8 pc=0x7b24bc
net/http.(*conn).serve(0x121a012bd440, {0xf6ad78, 0x1219ff547170})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:2135 +0x6bc fp=0x1219ff8c5fb8 sp=0x1219ff8c5b18 pc=0x764f3c
net/http.(*Server).Serve.gowrap3()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:3595 +0x1f fp=0x1219ff8c5fe0 sp=0x1219ff8c5fb8 pc=0x7a9e5f
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219ff8c5fe8 sp=0x1219ff8c5fe0 pc=0x48f641
created by net/http.(*Server).Serve in goroutine 8900
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/server.go:3595 +0x4cc
goroutine 2337 gp=0x1219ff034f00 m=nil [chan receive, 9 minutes]:
runtime.gopark(0x0?, 0x1219fed22dc8?, 0xd9?, 0x8b?, 0x1219fec64d50?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219fed22d38 sp=0x1219fed22d18 pc=0x487406
runtime.chanrecv(0x1219ff153e80, 0x0, 0x1)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/chan.go:659 +0x4bd fp=0x1219fed22db0 sp=0x1219fed22d38 pc=0x41763d
runtime.chanrecv1(0x1219fea402a0?, 0xa05460?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/chan.go:501 +0x12 fp=0x1219fed22dd8 sp=0x1219fed22db0 pc=0x417152
testing.tRunner.func1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2149 +0x425 fp=0x1219fed22f70 sp=0x1219fed22dd8 pc=0x523305
testing.tRunner(0x1219feff6908, 0xf6e448)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2206 +0x123 fp=0x1219fed22fc0 sp=0x1219fed22f70 pc=0x51dca3
testing.(*T).Run.gowrap1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2265 +0x1b fp=0x1219fed22fe0 sp=0x1219fed22fc0 pc=0x52395b
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219fed22fe8 sp=0x1219fed22fe0 pc=0x48f641
created by testing.(*T).Run in goroutine 1
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2265 +0x4d4
goroutine 4111 gp=0x1219fed35a40 m=nil [sync.WaitGroup.Wait, 9 minutes]:
runtime.gopark(0xff5c60?, 0x1?, 0x90?, 0x14?, 0xf27248?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:474 +0xc6 fp=0x1219ff197ca0 sp=0x1219ff197c80 pc=0x487406
runtime.goparkunlock(...)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:480
runtime.semacquire1(0x1219ff0eb518, 0x0, 0x1, 0x0, 0x19)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/sema.go:192 +0x249 fp=0x1219ff197d08 sp=0x1219ff197ca0 pc=0x462a89
sync.runtime_SemacquireWaitGroup(0x0?, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/sema.go:114 +0x2e fp=0x1219ff197d40 sp=0x1219ff197d08 pc=0x488b8e
sync.(*WaitGroup).Wait(0x1219ff0eb510)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/sync/waitgroup.go:206 +0x85 fp=0x1219ff197d68 sp=0x1219ff197d40 pc=0x497b45
net/http/httptest.(*Server).Close(0x1219ff0eb4a0)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/httptest/server.go:536 +0x28f fp=0x1219ff197e30 sp=0x1219ff197d68 pc=0x7c3d6f
net/http_test.(*clientServerTest).close(0x1219ff495400)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/clientserver_test.go:177 +0x2a fp=0x1219ff197e48 sp=0x1219ff197e30 pc=0x83906a
net/http_test.newClientServerTest.func2()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/net/http/clientserver_test.go:350 +0x17 fp=0x1219ff197e60 sp=0x1219ff197e48 pc=0x8cefd7
testing.(*common).Cleanup.func1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:1462 +0x109 fp=0x1219ff197ed0 sp=0x1219ff197e60 pc=0x522aa9
testing.(*common).runCleanup(0x1219ff0cc6c8, 0x0?)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:1822 +0xea fp=0x1219ff197f50 sp=0x1219ff197ed0 pc=0x51d26a
testing.tRunner.func2()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2194 +0x25 fp=0x1219ff197f70 sp=0x1219ff197f50 pc=0x522ec5
testing.tRunner(0x1219ff0cc6c8, 0x1219fec64d80)
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2206 +0x114 fp=0x1219ff197fc0 sp=0x1219ff197f70 pc=0x51dc94
testing.(*T).Run.gowrap1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2265 +0x1b fp=0x1219ff197fe0 sp=0x1219ff197fc0 pc=0x52395b
runtime.goexit({})
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/runtime/asm_amd64.s:1271 +0x1 fp=0x1219ff197fe8 sp=0x1219ff197fe0 pc=0x48f641
created by testing.(*T).Run in goroutine 2337
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2265 +0x4d4
FAIL net/http 600.080s
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /home/claude/golang_go_78438/go/src/net/http && sed -n '1195,1330p' transport_test.go
cd /home/claude/golang_go_78438/go/src/net/http && sed -n '1195,1330p' transport_test.go
Read testTransportGzip
}
if g, e := res.Header.Get("Content-Encoding"), test.accept; g != e {
t.Errorf("%d. Content-Encoding = %q; want %q", i, g, e)
}
}
}
func TestTransportGzip(t *testing.T) { run(t, testTransportGzip, http3SkippedMode) }
func testTransportGzip(t *testing.T, mode testMode) {
const testString = "The test string aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
const nRandBytes = 1024 * 1024
ts := newClientServerTest(t, mode, HandlerFunc(func(rw ResponseWriter, req *Request) {
if req.Method == "HEAD" {
if g := req.Header.Get("Accept-Encoding"); g != "" {
t.Errorf("HEAD request sent with Accept-Encoding of %q; want none", g)
}
return
}
if g, e := req.Header.Get("Accept-Encoding"), "gzip"; g != e {
t.Errorf("Accept-Encoding = %q, want %q", g, e)
}
rw.Header().Set("Content-Encoding", "gzip")
var w io.Writer = rw
var buf bytes.Buffer
if req.FormValue("chunked") == "0" {
w = &buf
defer io.Copy(rw, &buf)
defer func() {
rw.Header().Set("Content-Length", strconv.Itoa(buf.Len()))
}()
}
gz := gzip.NewWriter(w)
gz.Write([]byte(testString))
if req.FormValue("body") == "large" {
io.CopyN(gz, rand.Reader, nRandBytes)
}
gz.Close()
})).ts
c := ts.Client()
for _, chunked := range []string{"1", "0"} {
// First fetch something large, but only read some of it.
res, err := c.Get(ts.URL + "/?body=large&chunked=" + chunked)
if err != nil {
t.Fatalf("large get: %v", err)
}
buf := make([]byte, len(testString))
n, err := io.ReadFull(res.Body, buf)
if err != nil {
t.Fatalf("partial read of large response: size=%d, %v", n, err)
}
if e, g := testString, string(buf); e != g {
t.Errorf("partial read got %q, expected %q", g, e)
}
res.Body.Close()
// Read on the body, even though it's closed
n, err = res.Body.Read(buf)
if n != 0 || err == nil {
t.Errorf("expected error post-closed large Read; got = %d, %v", n, err)
}
// Then something small.
res, err = c.Get(ts.URL + "/?chunked=" + chunked)
if err != nil {
t.Fatal(err)
}
body, err := io.ReadAll(res.Body)
if err != nil {
t.Fatal(err)
}
if g, e := string(body), testString; g != e {
t.Fatalf("body = %q; want %q", g, e)
}
if g, e := res.Header.Get("Content-Encoding"), ""; g != e {
t.Fatalf("Content-Encoding = %q; want %q", g, e)
}
// Read on the body after it's been fully read:
n, err = res.Body.Read(buf)
if n != 0 || err == nil {
t.Errorf("expected Read error after exhausted reads; got %d, %v", n, err)
}
res.Body.Close()
n, err = res.Body.Read(buf)
if n != 0 || err == nil {
t.Errorf("expected Read error after Close; got %d, %v", n, err)
}
}
// And a HEAD request too, because they're always weird.
res, err := c.Head(ts.URL)
if err != nil {
t.Fatalf("Head: %v", err)
}
if res.StatusCode != 200 {
t.Errorf("Head status=%d; want=200", res.StatusCode)
}
}
// A transport100Continue test exercises Transport behaviors when sending a
// request with an Expect: 100-continue header.
type transport100ContinueTest struct {
t *testing.T
reqdone chan struct{}
resp *Response
respErr error
conn net.Conn
reader *bufio.Reader
}
const transport100ContinueTestBody = "request body"
// newTransport100ContinueTest creates a Transport and sends an Expect: 100-continue
// request on it.
func newTransport100ContinueTest(t *testing.T, timeout time.Duration) *transport100ContinueTest {
ln := newLocalListener(t)
defer ln.Close()
test := &transport100ContinueTest{
t: t,
reqdone: make(chan struct{}),
}
tr := &Transport{
ExpectContinueTimeout: timeout,
}
go func() {
defer close(test.reqdone)
body := strings.NewReader(transport100ContinueTestBody)
req, _ := NewRequest("PUT", "http://"+ln.Addr().String(), body)
req.Header.Set("Expect", "100-continue")
req.ContentLength = int64(len(transport100ContinueTestBody))Now let me look at the Solaris netpoll implementation and check the other failure logs.
Ran tools (35 Bash)
Bash cd /home/claude/golang_go_78438/go/src/runtime && sed -n '1,120p' netpoll_solaris.go
cd /home/claude/golang_go_78438/go/src/runtime && sed -n '1,120p' netpoll_solaris.go
Read netpoll_solaris.go part 1
// Copyright 2014 The Go Authors. All rights reserved.
// Use of this source code is governed by a BSD-style
// license that can be found in the LICENSE file.
package runtime
import (
"internal/goarch"
"internal/runtime/atomic"
"unsafe"
)
// Solaris runtime-integrated network poller.
//
// Solaris uses event ports for scalable network I/O. Event
// ports are level-triggered, unlike epoll and kqueue which
// can be configured in both level-triggered and edge-triggered
// mode. Level triggering means we have to keep track of a few things
// ourselves. After we receive an event for a file descriptor,
// it's our responsibility to ask again to be notified for future
// events for that descriptor. When doing this we must keep track of
// what kind of events the goroutines are currently interested in,
// for example a fd may be open both for reading and writing.
//
// A description of the high level operation of this code
// follows. Networking code will get a file descriptor by some means
// and will register it with the netpolling mechanism by a code path
// that eventually calls runtime·netpollopen. runtime·netpollopen
// calls port_associate with an empty event set. That means that we
// will not receive any events at this point. The association needs
// to be done at this early point because we need to process the I/O
// readiness notification at some point in the future. If I/O becomes
// ready when nobody is listening, when we finally care about it,
// nobody will tell us anymore.
//
// Beside calling runtime·netpollopen, the networking code paths
// will call runtime·netpollarm each time goroutines are interested
// in doing network I/O. Because now we know what kind of I/O we
// are interested in (reading/writing), we can call port_associate
// passing the correct type of event set (POLLIN/POLLOUT). As we made
// sure to have already associated the file descriptor with the port,
// when we now call port_associate, we will unblock the main poller
// loop (in runtime·netpoll) right away if the socket is actually
// ready for I/O.
//
// The main poller loop runs in its own thread waiting for events
// using port_getn. When an event happens, it will tell the scheduler
// about it using runtime·netpollready. Besides doing this, it must
// also re-associate the events that were not part of this current
// notification with the file descriptor. Failing to do this would
// mean each notification will prevent concurrent code using the
// same file descriptor in parallel.
//
// The logic dealing with re-associations is encapsulated in
// runtime·netpollupdate. This function takes care to associate the
// descriptor only with the subset of events that were previously
// part of the association, except the one that just happened. We
// can't re-associate with that right away, because event ports
// are level triggered so it would cause a busy loop. Instead, that
// association is effected only by the runtime·netpollarm code path,
// when Go code actually asks for I/O.
//
// The open and arming mechanisms are serialized using the lock
// inside PollDesc. This is required because the netpoll loop runs
// asynchronously in respect to other Go code and by the time we get
// to call port_associate to update the association in the loop, the
// file descriptor might have been closed and reopened already. The
// lock allows runtime·netpollupdate to be called synchronously from
// the loop thread while preventing other threads operating to the
// same PollDesc, so once we unblock in the main loop, until we loop
// again we know for sure we are always talking about the same file
// descriptor and can safely access the data we want (the event set).
//go:cgo_import_dynamic libc_port_create port_create "libc.so"
//go:cgo_import_dynamic libc_port_associate port_associate "libc.so"
//go:cgo_import_dynamic libc_port_dissociate port_dissociate "libc.so"
//go:cgo_import_dynamic libc_port_getn port_getn "libc.so"
//go:cgo_import_dynamic libc_port_alert port_alert "libc.so"
//go:linkname libc_port_create libc_port_create
//go:linkname libc_port_associate libc_port_associate
//go:linkname libc_port_dissociate libc_port_dissociate
//go:linkname libc_port_getn libc_port_getn
//go:linkname libc_port_alert libc_port_alert
var (
libc_port_create,
libc_port_associate,
libc_port_dissociate,
libc_port_getn,
libc_port_alert libcFunc
netpollWakeSig atomic.Uint32 // used to avoid duplicate calls of netpollBreak
)
func errno() int32 {
return *getg().m.perrno
}
func port_create() int32 {
return int32(sysvicall0(&libc_port_create))
}
func port_associate(port, source int32, object uintptr, events uint32, user uintptr) int32 {
return int32(sysvicall5(&libc_port_associate, uintptr(port), uintptr(source), object, uintptr(events), user))
}
func port_dissociate(port, source int32, object uintptr) int32 {
return int32(sysvicall3(&libc_port_dissociate, uintptr(port), uintptr(source), object))
}
func port_getn(port int32, evs *portevent, max uint32, nget *uint32, timeout *timespec) int32 {
return int32(sysvicall5(&libc_port_getn, uintptr(port), uintptr(unsafe.Pointer(evs)), uintptr(max), uintptr(unsafe.Pointer(nget)), uintptr(unsafe.Pointer(timeout))))
}
func port_alert(port int32, flags, events uint32, user uintptr) int32 {
return int32(sysvicall4(&libc_port_alert, uintptr(port), uintptr(flags), uintptr(events), user))
}
var portfd int32 = -1Bash cd /tmp; for id in 8681423021638815105 8676093236579846177; do echo "=== $id"; curl -s -X POST "https://cr-buildbucket.a...
cd /tmp; for id in 8681423021638815105 8676093236579846177; do echo "=== $id"; curl -s -X POST "https://cr-buildbucket.appspot.com/prpc/buildbucket.v2.Builds/GetBuild" -H "Content-Type: application/json" -d "{\"id\":\"$id\",\"mask\":{\"fields\":\"steps\"}}" -o b_$id.json; python3 -c "
import json,sys
s=open('b_$id.json').read(); s=s[s.index('{'):]
d=json.loads(s)
for st in d.get('steps',[]):
if st.get('status')!='SUCCESS':
for l in st.get('logs',[]):
print(st['name'],'|',l['name'],'|',l.get('viewUrl'))
"; doneGet logs for other builds
Exit code 1 === 8681423021638815105 Traceback (most recent call last): File "<string>", line 3, in <module> File "<frozen codecs>", line 322, in decode UnicodeDecodeError: 'utf-8' codec can't decode byte 0x8a in position 0: invalid start byte === 8676093236579846177 Traceback (most recent call last): File "<string>", line 3, in <module> File "<frozen codecs>", line 322, in decode UnicodeDecodeError: 'utf-8' codec can't decode byte 0x8a in position 0: invalid start byte
Bash cd /tmp; for id in 8681423021638815105 8676093236579846177 8681022390813019185 8675856801014264801 8686112892523717025; ...
cd /tmp; for id in 8681423021638815105 8676093236579846177 8681022390813019185 8675856801014264801 8686112892523717025; do echo "=== $id"; curl -s -X POST "https://cr-buildbucket.appspot.com/prpc/buildbucket.v2.Builds/GetBuild" -H "Content-Type: application/json" -H "Accept: application/json" -d "{\"id\":\"$id\",\"mask\":{\"fields\":\"builder,summaryMarkdown\"}}" | tail -c +6 | python3 -c "
import json,sys
d=json.load(sys.stdin)
print(d['builder']['builder'])
print(d['summaryMarkdown'][:800])
"; doneGet summary for all builds
=== 8681423021638815105 gotip-linux-ppc64le_power8 * go on master ([commit 2e67b18](https://go.googlesource.com/go/+/2e67b18935b9ddfaf36269bf2fd53a364e7a94d9)) To reproduce, try `gomote repro 8681423021638815105`. Additional links for debugging: * [go tool dist test -json output](https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8681423021638815105/+/u/step/11/log/2) === 8676093236579846177 go1.27-solaris-amd64 * go on release-branch.go1.27 ([commit 96307fb](https://go.googlesource.com/go/+/96307fb52ff61e791cec8d0578237c450f5c4eea)) To reproduce, try `gomote repro 8676093236579846177`. Additional links for debugging: * [go tool dist test -json output](https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8676093236579846177/+/u/step/11/log/2) === 8681022390813019185 gotip-solaris-amd64 * go on master ([commit 05f75fb](https://go.googlesource.com/go/+/05f75fb9e863500d2c24143becfff6b06670c559)) To reproduce, try `gomote repro 8681022390813019185`. Additional links for debugging: * [go tool dist test -json output](https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8681022390813019185/+/u/step/11/log/2) === 8675856801014264801 gotip-solaris-amd64 * go on master ([commit c29765a](https://go.googlesource.com/go/+/c29765a803e2164b17bc669fe944ec23fa98bef0)) To reproduce, try `gomote repro 8675856801014264801`. Additional links for debugging: * [go tool dist test -json output](https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8675856801014264801/+/u/step/11/log/2) === 8686112892523717025 gotip-solaris-amd64 * go on master ([commit 1fd6879](https://go.googlesource.com/go/+/1fd68799c39bd4a3f7e16a1ee24fcaca3efe5357)) To reproduce, try `gomote repro 8686112892523717025`. Additional links for debugging: * [go tool dist test -json output](https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8686112892523717025/+/u/step/11/log/2) Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8681423021638815105/+/u/step/11/log/...
cd /tmp; curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8681423021638815105/+/u/step/11/log/3?format=raw" 2>/dev/null | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('ppc.txt','w').write(''.join(out))
print('done')
"; grep -n "panic: test timed out" -A6 ppc.txt | head -20Extract ppc64le net/http output
done
5107:panic: test timed out after 18m0s
5108- running tests:
5109- TestTransportGzip/h1 (17m57s)
5110-
5111-goroutine 16921 gp=0x1917eda00700 m=9 mp=0x1917ed8b8008 [running]:
5112-panic({0xa7e0f8?, 0x1917ed7a78b0?})
5113- /home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/panic.go:878 +0x188 fp=0x1917ed990ee0 sp=0x1917ed990e28 pc=0xa4678
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; grep -n "blocked in Close" ppc.txt | head; echo "=== stuck goroutine frames:"; grep -n "net/http_test\.testTran...
cd /tmp; grep -n "blocked in Close" ppc.txt | head; echo "=== stuck goroutine frames:"; grep -n "net/http_test\.testTransportGzip\|httptest.(\*Server).Close\|internal/poll.(\*FD).Write\|internal/poll.(\*FD).Read" ppc.txt | head -20
Compare ppc64le signature
=== stuck goroutine frames:
5322:net/http_test.testTransportGzip(0x1917edd1f688, {0x5900a0, 0x2})
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; sed -n '5280,5345p' ppc.txt
cd /tmp; sed -n '5280,5345p' ppc.txt
Show ppc64le stuck test goroutine
created by runtime.gcBgMarkStartWorkers in goroutine 11
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/mgc.go:1711 +0x164
goroutine 2095 gp=0x1917eda01340 m=nil [chan receive, 17 minutes]:
runtime.gopark(0x5e7700?, 0xa7325202045?, 0x0?, 0x0?, 0x2a680?)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:467 +0x118 fp=0x1917ee0bfcb8 sp=0x1917ee0bfc80 pc=0xa4ba8
runtime.chanrecv(0x1917edfd4180, 0x0, 0x1)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/chan.go:667 +0x560 fp=0x1917ee0bfd58 sp=0x1917ee0bfcb8 pc=0x2ad70
runtime.chanrecv1(0x1917ee0c2310?, 0x5ed4c0?)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/chan.go:509 +0x10 fp=0x1917ee0bfd90 sp=0x1917ee0bfd58 pc=0x2a7d0
testing.tRunner.func1()
/home/swarming/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2147 +0x3fc fp=0x1917ee0bff30 sp=0x1917ee0bfd90 pc=0x154c7c
testing.tRunner(0x1917ee0b8fc8, 0xb00888)
/home/swarming/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2196 +0x14c fp=0x1917ee0bff90 sp=0x1917ee0bff30 pc=0x14f25c
testing.(*T).Run.gowrap1()
/home/swarming/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2255 +0x2c fp=0x1917ee0bffc0 sp=0x1917ee0bff90 pc=0x15529c
runtime.goexit({})
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/asm_ppc64x.s:1057 +0x4 fp=0x1917ee0bffc0 sp=0x1917ee0bffc0 pc=0xad174
created by testing.(*T).Run in goroutine 1
/home/swarming/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2255 +0x488
goroutine 3194 gp=0x1917ee450a80 m=nil [select, 17 minutes]:
runtime.gopark(0x1917edefd5c0?, 0x2?, 0x1c?, 0x4b?, 0x1917edefd4f4?)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:467 +0x118 fp=0x1917edefd330 sp=0x1917edefd2f8 pc=0xa4ba8
runtime.selectgo(0x1917edefd5c0, 0x1917edefd4f0, 0xf?, 0x0, 0xf?, 0x1)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/select.go:351 +0xa34 fp=0x1917edefd498 sp=0x1917edefd330 pc=0x7aef4
net/http.(*Transport).getConn(0x1917ee018c40, 0x1917eead7770, {{}, 0x0, {0x1917eda8f3b0, 0x4}, {0x1917eeba4860, 0xf}, 0x0})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/transport.go:1619 +0x420 fp=0x1917edefd680 sp=0x1917edefd498 pc=0x394bb0
net/http.(*Transport).roundTrip(0x1917ee018c40, 0x1917edaf7400)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/transport.go:714 +0x960 fp=0x1917edefd838 sp=0x1917edefd680 pc=0x390570
net/http.(*Transport).RoundTrip(0x7839025c2401?, 0xaf8b40?)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/roundtrip.go:33 +0x3c fp=0x1917edefd868 sp=0x1917edefd838 pc=0x3ca0cc
net/http.send(0x1917edaf7400, {0xaf8b40, 0x1917ee018c40}, {0x1917eeaef0e0?, 0x1e2cb8?, 0x0?})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/client.go:264 +0x580 fp=0x1917edefda40 sp=0x1917edefd868 pc=0x35a2a0
net/http.(*Client).send(0x1917eeac7b00, 0x1917edaf7400, {0x1917ed922b70?, 0x33d8c?, 0x0?})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/client.go:185 +0x268 fp=0x1917edefdae0 sp=0x1917edefda40 pc=0x359bc8
net/http.(*Client).do(0x1917eeac7b00, 0x1917edaf7400)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/client.go:736 +0x898 fp=0x1917edefdd08 sp=0x1917edefdae0 pc=0x35ba58
net/http.(*Client).Do(...)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/client.go:595
net/http.(*Client).Get(0x1917eeac7b00, {0x1917eda8f3b0?, 0x3b?})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/client.go:489 +0x9c fp=0x1917edefdd68 sp=0x1917edefdd08 pc=0x35b0cc
net/http_test.testTransportGzip(0x1917edd1f688, {0x5900a0, 0x2})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/transport_test.go:1258 +0x414 fp=0x1917edefdeb8 sp=0x1917edefdd68 pc=0x4a6f84
net/http_test.run[...].func1()
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/clientserver_test.go:130 +0x124 fp=0x1917edefdf30 sp=0x1917edefdeb8 pc=0x524fb4
testing.tRunner(0x1917edd1f688, 0x1917ed8ce9c0)
/home/swarming/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2190 +0x108 fp=0x1917edefdf90 sp=0x1917edefdf30 pc=0x14f218
testing.(*T).Run.gowrap1()
/home/swarming/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2255 +0x2c fp=0x1917edefdfc0 sp=0x1917edefdf90 pc=0x15529c
runtime.goexit({})
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/asm_ppc64x.s:1057 +0x4 fp=0x1917edefdfc0 sp=0x1917edefdfc0 pc=0xad174
created by testing.(*T).Run in goroutine 2095
/home/swarming/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2255 +0x488
goroutine 7203 gp=0x1917ee59afc0 m=nil [IO wait, 17 minutes]:
runtime.gopark(0x0?, 0x0?, 0x0?, 0x0?, 0x0?)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:467 +0x118 fp=0x1917eddadaf0 sp=0x1917eddadab8 pc=0xa4ba8
runtime.netpollblock(0x0?, 0x0?, 0x0?)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/netpoll.go:575 +0x188 fp=0x1917eddadb38 sp=0x1917eddadaf0 pc=0x5ef68
internal/poll.runtime_pollWait(0x78389bf65800, 0x72)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/netpoll.go:351 +0xf4 fp=0x1917eddadb78 sp=0x1917eddadb38 pc=0xa3c24
internal/poll.(*pollDesc).wait(0x1917ed83d580?, 0x3011c?, 0x0)
/home/swarming/.swarming/w/ir/x/w/goroot/src/internal/poll/fd_poll_runtime.go:84 +0x4c fp=0x1917eddadbd0 sp=0x1917eddadb78 pc=0xe500c
internal/poll.(*pollDesc).waitRead(...)
/home/swarming/.swarming/w/ir/x/w/goroot/src/internal/poll/fd_poll_runtime.go:89
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; python3 - <<'EOF' ...
cd /tmp; python3 - <<'EOF'
import re
s=open('ppc.txt').read()
i=s.index('panic: test timed out')
blocks=re.split(r'\n\n', s[i:])
print('num blocks', len(blocks))
for b in blocks:
lines=b.strip().split('\n')
if not lines or not lines[0].startswith('goroutine'): continue
hdr=lines[0]
frames=[l for l in lines if not l.startswith('\t') and not l.startswith('created by')][1:]
# keep interesting
interesting=[f.split('(')[0] for f in frames]
print(hdr)
print(' ', ' <- '.join(interesting[:8]))
EOFSummarize ppc64le goroutines
num blocks 18
goroutine 16921 gp=0x1917eda00700 m=9 mp=0x1917ed8b8008 [running]:
panic <- testing. <- runtime.goexit
goroutine 1 gp=0x1917ed7761c0 m=nil [chan receive, 17 minutes]:
runtime.gopark <- runtime.chanrecv <- runtime.chanrecv1 <- testing.tRunner.func1 <- testing.tRunner <- testing.runTests <- testing. <- net/http_test.TestMain
goroutine 2 gp=0x1917ed776c40 m=nil [force gc (idle), 3 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.forcegchelper <- runtime.goexit
goroutine 3 gp=0x1917ed777180 m=nil [GC sweep wait]:
runtime.gopark <- runtime.goparkunlock <- runtime.bgsweep <- runtime.gcenable.gowrap1 <- runtime.goexit
goroutine 4 gp=0x1917ed777340 m=nil [GC scavenge wait]:
runtime.gopark <- runtime.goparkunlock <- runtime. <- runtime.bgscavenge <- runtime.gcenable.gowrap2 <- runtime.goexit
goroutine 5 gp=0x1917ed777880 m=nil [GOMAXPROCS updater (idle), 17 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.updateMaxProcsGoroutine <- runtime.goexit
goroutine 6 gp=0x1917ed820000 m=nil [finalizer wait, 17 minutes]:
runtime.gopark <- runtime.runFinalizers <- runtime.goexit
goroutine 7 gp=0x1917ed8201c0 m=nil [cleanup wait, 17 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime. <- runtime.runCleanups <- runtime.goexit
goroutine 12 gp=0x1917ed820540 m=nil [GC worker (idle), 17 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 13 gp=0x1917ed820700 m=nil [GC worker (idle), 17 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 14 gp=0x1917ed8208c0 m=nil [GC worker (idle), 17 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 15 gp=0x1917ed820a80 m=nil [GC worker (idle)]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 16 gp=0x1917ed820c40 m=nil [GC worker (idle), 17 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 2095 gp=0x1917eda01340 m=nil [chan receive, 17 minutes]:
runtime.gopark <- runtime.chanrecv <- runtime.chanrecv1 <- testing.tRunner.func1 <- testing.tRunner <- testing. <- runtime.goexit
goroutine 3194 gp=0x1917ee450a80 m=nil [select, 17 minutes]:
runtime.gopark <- runtime.selectgo <- net/http. <- net/http. <- net/http. <- net/http.send <- net/http. <- net/http.
goroutine 7203 gp=0x1917ee59afc0 m=nil [IO wait, 17 minutes]:
runtime.gopark <- runtime.netpollblock <- internal/poll.runtime_pollWait <- internal/poll. <- internal/poll. <- internal/poll. <- net. <- net.
goroutine 7213 gp=0x1917eda9cc40 m=nil [runnable]:
runtime.asyncPreempt2 <- runtime.asyncPreempt <- context. <- net. <- net. <- net. <- net/http/httptest. <- net/http.
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; python3 - <<'EOF' ...
cd /tmp; python3 - <<'EOF'
import re
s=open('ppc.txt').read()
i=s.index('panic: test timed out')
blocks=re.split(r'\n\n', s[i:])
for b in blocks:
if b.startswith('goroutine 7203') or b.startswith('goroutine 7213'):
print(b)
print('========')
EOFShow ppc64le net goroutines full stacks
goroutine 7203 gp=0x1917ee59afc0 m=nil [IO wait, 17 minutes]:
runtime.gopark(0x0?, 0x0?, 0x0?, 0x0?, 0x0?)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/proc.go:467 +0x118 fp=0x1917eddadaf0 sp=0x1917eddadab8 pc=0xa4ba8
runtime.netpollblock(0x0?, 0x0?, 0x0?)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/netpoll.go:575 +0x188 fp=0x1917eddadb38 sp=0x1917eddadaf0 pc=0x5ef68
internal/poll.runtime_pollWait(0x78389bf65800, 0x72)
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/netpoll.go:351 +0xf4 fp=0x1917eddadb78 sp=0x1917eddadb38 pc=0xa3c24
internal/poll.(*pollDesc).wait(0x1917ed83d580?, 0x3011c?, 0x0)
/home/swarming/.swarming/w/ir/x/w/goroot/src/internal/poll/fd_poll_runtime.go:84 +0x4c fp=0x1917eddadbd0 sp=0x1917eddadb78 pc=0xe500c
internal/poll.(*pollDesc).waitRead(...)
/home/swarming/.swarming/w/ir/x/w/goroot/src/internal/poll/fd_poll_runtime.go:89
internal/poll.(*FD).Accept(0x1917ed83d580)
/home/swarming/.swarming/w/ir/x/w/goroot/src/internal/poll/fd_unix.go:613 +0x280 fp=0x1917eddadc88 sp=0x1917eddadbd0 pc=0xe7670
net.(*netFD).accept(0x1917ed83d580)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/fd_unix.go:149 +0x38 fp=0x1917eddadd60 sp=0x1917eddadc88 pc=0x24de78
net.(*TCPListener).accept(0x1917eea98000)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/tcpsock_posix.go:159 +0x34 fp=0x1917eddaddc0 sp=0x1917eddadd60 pc=0x260644
net.(*TCPListener).Accept(0x1917eea98000)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/tcpsock.go:387 +0x44 fp=0x1917eddade20 sp=0x1917eddaddc0 pc=0x25f894
net/http.(*Server).Serve(0x1917ede4a6c0, {0xafca30, 0x1917eea98000})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/server.go:3523 +0x424 fp=0x1917eddadf68 sp=0x1917eddade20 pc=0x384204
net/http/httptest.(*Server).goServe.func1()
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/httptest/server.go:352 +0x64 fp=0x1917eddadfc0 sp=0x1917eddadf68 pc=0x3d88a4
runtime.goexit({})
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/asm_ppc64x.s:1057 +0x4 fp=0x1917eddadfc0 sp=0x1917eddadfc0 pc=0xad174
created by net/http/httptest.(*Server).goServe in goroutine 3194
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/httptest/server.go:350 +0x80
========
goroutine 7213 gp=0x1917eda9cc40 m=nil [runnable]:
runtime.asyncPreempt2()
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/preempt.go:320 +0x14 fp=0x1917edcdb3f0 sp=0x1917edcdb3c8 pc=0x64b94
runtime.asyncPreempt()
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/preempt_ppc64x.s:75 +0x10c fp=0x1917edcdb5f8 sp=0x1917edcdb3f0 pc=0xadc4c
context.(*cancelCtx).Deadline(0x6a0ba156?)
<autogenerated>:1 +0x10 fp=0x1917edcdb618 sp=0x1917edcdb618 pc=0x119400
net.(*Dialer).deadline(0x1917ee27f710, {0xafd4a0?, 0x1917eead77c0?}, {0xae85f8?, 0x1917eeac66f0?, 0xb89fe0?})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/dial.go:258 +0xc8 fp=0x1917edcdb680 sp=0x1917edcdb618 pc=0x2432f8
net.(*Dialer).dialCtx(0x1917ee27f710, {0xafd4a0, 0x1917eead77c0})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/dial.go:567 +0x78 fp=0x1917edcdb728 sp=0x1917edcdb680 pc=0x244d48
net.(*Dialer).DialContext(0x1917ee27f710, {0xafd4a0?, 0x1917eead77c0?}, {0x590282, 0x3}, {0x1917eeba4860, 0xf})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/dial.go:530 +0x94 fp=0x1917edcdb870 sp=0x1917edcdb728 pc=0x2446e4
net/http/httptest.(*Server).Start.func1({0xafd4a0, 0x1917eead77c0}, {0x590282, 0x3}, {0x1917eeba4860?, 0xf?})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/httptest/server.go:153 +0x188 fp=0x1917edcdb8d8 sp=0x1917edcdb870 pc=0x3d8068
net/http.(*Transport).dial(0x1?, {0xafd4a0?, 0x1917eead77c0?}, {0x590282?, 0x1917eda8f3b0?}, {0x1917eeba4860?, 0x1917eeba4860?})
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/transport.go:1374 +0x13c fp=0x1917edcdb950 sp=0x1917edcdb8d8 pc=0x393d1c
net/http.(*Transport).dialConn(0x1917ee018c40, {0xafd4a0, 0x1917eead77c0}, {{}, 0x0, {0x1917eda8f3b0, 0x4}, {0x1917eeba4860, 0xf}, 0x0}, ...)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/transport.go:1920 +0x820 fp=0x1917edcdbec8 sp=0x1917edcdb950 pc=0x3966c0
net/http.(*Transport).dialConnFor(0x1917ee018c40, 0x1917edb3f080)
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/transport.go:1715 +0x98 fp=0x1917edcdbf60 sp=0x1917edcdbec8 pc=0x395418
net/http.(*Transport).startDialConnForLocked.func1()
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/transport.go:1696 +0x3c fp=0x1917edcdbfc0 sp=0x1917edcdbf60 pc=0x3c1f6c
runtime.goexit({})
/home/swarming/.swarming/w/ir/x/w/goroot/src/runtime/asm_ppc64x.s:1057 +0x4 fp=0x1917edcdbfc0 sp=0x1917edcdbfc0 pc=0xad174
created by net/http.(*Transport).startDialConnForLocked in goroutine 3194
/home/swarming/.swarming/w/ir/x/w/goroot/src/net/http/transport.go:1695 +0x148
FAIL net/http 1080.034s
========
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /home/claude/golang_go_78438/go/src/net/http && sed -n '1590,1730p' transport.go
cd /home/claude/golang_go_78438/go/src/net/http && sed -n '1590,1730p' transport.go
Read getConn/queueForDial
}
// getConn dials and creates a new persistConn to the target as
// specified in the connectMethod. This includes doing a proxy CONNECT
// and/or setting up TLS. If this doesn't return an error, the persistConn
// is ready to write requests to.
func (t *Transport) getConn(treq *transportRequest, cm connectMethod) (_ *persistConn, err error) {
req := treq.Request
trace := treq.trace
ctx := req.Context()
if trace != nil && trace.GetConn != nil {
trace.GetConn(cm.addr())
}
// Detach from the request context's cancellation signal.
// The dial should proceed even if the request is canceled,
// because a future request may be able to make use of the connection.
//
// We retain the request context's values.
dialCtx, dialCancel := context.WithCancel(context.WithoutCancel(ctx))
w := &wantConn{
cm: cm,
key: cm.key(),
ctx: dialCtx,
cancelCtx: dialCancel,
result: make(chan connOrError, 1),
beforeDial: testHookPrePendingDial,
afterDial: testHookPostPendingDial,
}
defer func() {
if err != nil {
w.cancel(t)
}
}()
// Queue for idle connection.
if delivered := t.queueForIdleConn(w); !delivered {
t.queueForDial(w)
}
// Wait for completion or cancellation.
select {
case r := <-w.result:
// Trace success but only for HTTP/1.
// HTTP/2 calls trace.GotConn itself.
if r.pc != nil && r.pc.alt == nil && trace != nil && trace.GotConn != nil {
info := httptrace.GotConnInfo{
Conn: r.pc.conn,
Reused: r.pc.isReused(),
}
if !r.idleAt.IsZero() {
info.WasIdle = true
info.IdleTime = time.Since(r.idleAt)
}
trace.GotConn(info)
}
if r.err != nil {
// If the request has been canceled, that's probably
// what caused r.err; if so, prefer to return the
// cancellation error (see golang.org/issue/16049).
select {
case <-treq.ctx.Done():
err := context.Cause(treq.ctx)
if err == errRequestCanceled {
err = errRequestCanceledConn
}
return nil, err
default:
// return below
}
}
return r.pc, r.err
case <-treq.ctx.Done():
err := context.Cause(treq.ctx)
if err == errRequestCanceled {
err = errRequestCanceledConn
}
return nil, err
}
}
// queueForDial queues w to wait for permission to begin dialing.
// Once w receives permission to dial, it will do so in a separate goroutine.
func (t *Transport) queueForDial(w *wantConn) {
w.beforeDial()
t.connsPerHostMu.Lock()
defer t.connsPerHostMu.Unlock()
if t.MaxConnsPerHost <= 0 {
t.startDialConnForLocked(w)
return
}
if n := t.connsPerHost[w.key]; n < t.MaxConnsPerHost {
if t.connsPerHost == nil {
t.connsPerHost = make(map[connectMethodKey]int)
}
t.connsPerHost[w.key] = n + 1
t.startDialConnForLocked(w)
return
}
if t.connsPerHostWait == nil {
t.connsPerHostWait = make(map[connectMethodKey]wantConnQueue)
}
q := t.connsPerHostWait[w.key]
q.cleanFrontNotWaiting()
q.pushBack(w)
t.connsPerHostWait[w.key] = q
}
// startDialConnFor calls dialConn in a new goroutine.
// t.connsPerHostMu must be held.
func (t *Transport) startDialConnForLocked(w *wantConn) {
t.dialsInProgress.cleanFrontCanceled()
t.dialsInProgress.pushBack(w)
go func() {
t.dialConnFor(w)
t.connsPerHostMu.Lock()
defer t.connsPerHostMu.Unlock()
w.cancelCtx = nil
}()
}
// dialConnFor dials on behalf of w and delivers the result to w.
// dialConnFor has received permission to dial w.cm and is counted in t.connCount[w.cm.key()].
// If the dial is canceled or unsuccessful, dialConnFor decrements t.connCount[w.cm.key()].
func (t *Transport) dialConnFor(w *wantConn) {
defer w.afterDial()
ctx := w.getCtxForDial()
if ctx == nil {
t.decConnsPerHost(w.key)
return
}
const isClientConn = false
pc, err := t.dialConn(ctx, w.cm, isClientConn, nil)
delivered := w.tryDeliver(pc, err, time.Time{})
if err == nil && (!delivered || pc.alt != nil) {Bash cd /tmp; python3 - <<'EOF' ...
cd /tmp; python3 - <<'EOF'
import re
s=open('nethttp.txt').read()
i=s.index('panic: test timed out')
blocks=re.split(r'\n\n', s[i:])
print('blocks',len(blocks))
for b in blocks:
lines=b.strip().split('\n')
if not lines or not lines[0].startswith('goroutine'): continue
frames=[l.split('(')[0] for l in lines if not l.startswith('\t') and not l.startswith('created by')][1:]
print(lines[0])
print(' ', ' <- '.join(frames[:10]))
EOFSummarize solaris goroutines
blocks 21
goroutine 19549 gp=0x1219ff816000 m=33 mp=0x1219feb4d008 [running]:
panic <- testing. <- runtime.goexit
goroutine 1 gp=0x1219fe8ac1e0 m=nil [chan receive, 9 minutes]:
runtime.gopark <- runtime.chanrecv <- runtime.chanrecv1 <- testing.tRunner.func1 <- testing.tRunner <- testing.runTests <- testing. <- net/http_test.TestMain <- main.main <- runtime.main
goroutine 2 gp=0x1219fe8acd20 m=nil [force gc (idle), 3 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.forcegchelper <- runtime.goexit
goroutine 3 gp=0x1219fe8ad2c0 m=nil [GC sweep wait]:
runtime.gopark <- runtime.goparkunlock <- runtime.bgsweep <- runtime.gcenable.gowrap1 <- runtime.goexit
goroutine 4 gp=0x1219fe8ad4a0 m=nil [GC scavenge wait]:
runtime.gopark <- runtime.goparkunlock <- runtime. <- runtime.bgscavenge <- runtime.gcenable.gowrap2 <- runtime.goexit
goroutine 5 gp=0x1219fe8ada40 m=nil [GOMAXPROCS updater (idle), 10 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.updateMaxProcsGoroutine <- runtime.goexit
goroutine 6 gp=0x1219fe9501e0 m=nil [finalizer wait, 9 minutes]:
runtime.gopark <- runtime.runFinalizers <- runtime.goexit
goroutine 18 gp=0x1219fe9da000 m=nil [cleanup wait, 9 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime. <- runtime.runCleanups <- runtime.goexit
goroutine 7 gp=0x1219fe950780 m=nil [GC worker (idle), 9 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 8 gp=0x1219fe950960 m=nil [GC worker (idle), 9 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 9 gp=0x1219fe950b40 m=nil [GC worker (idle), 9 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 10 gp=0x1219fe950d20 m=nil [GC worker (idle), 9 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 11 gp=0x1219fe950f00 m=nil [GC worker (idle), 3 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 12 gp=0x1219fe9510e0 m=nil [GC worker (idle), 9 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 13 gp=0x1219fe9512c0 m=nil [GC worker (idle), 3 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 14 gp=0x1219fe9514a0 m=nil [GC worker (idle), 9 minutes]:
runtime.gopark <- runtime.gcBgMarkWorker <- runtime.gcBgMarkStartWorkers.gowrap1 <- runtime.goexit
goroutine 903 gp=0x1219fee6a5a0 m=nil [cleanup wait, 9 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime. <- runtime.runCleanups <- runtime.goexit
goroutine 8885 gp=0x1219feb10d20 m=nil [IO wait, 9 minutes]:
runtime.gopark <- runtime.netpollblock <- internal/poll.runtime_pollWait <- internal/poll. <- internal/poll. <- internal/poll. <- net. <- net. <- net/http.checkConnErrorWriter.Write <- bufio.
goroutine 2337 gp=0x1219ff034f00 m=nil [chan receive, 9 minutes]:
runtime.gopark <- runtime.chanrecv <- runtime.chanrecv1 <- testing.tRunner.func1 <- testing.tRunner <- testing. <- runtime.goexit
goroutine 4111 gp=0x1219fed35a40 m=nil [sync.WaitGroup.Wait, 9 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.semacquire1 <- sync.runtime_SemacquireWaitGroup <- sync. <- net/http/httptest. <- net/http_test. <- net/http_test.newClientServerTest.func2 <- testing. <- testing.
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /home/claude/golang_go_78438/go/src/runtime && sed -n '120,320p' netpoll_solaris.go
cd /home/claude/golang_go_78438/go/src/runtime && sed -n '120,320p' netpoll_solaris.go
Read netpoll_solaris core
func netpollinit() {
portfd = port_create()
if portfd >= 0 {
closeonexec(portfd)
return
}
print("runtime: port_create failed (errno=", errno(), ")\n")
throw("runtime: netpollinit failed")
}
func netpollIsPollDescriptor(fd uintptr) bool {
return fd == uintptr(portfd)
}
func netpollopen(fd uintptr, pd *pollDesc) int32 {
lock(&pd.lock)
// We don't register for any specific type of events yet, that's
// netpollarm's job. We merely ensure we call port_associate before
// asynchronous connect/accept completes, so when we actually want
// to do any I/O, the call to port_associate (from netpollarm,
// with the interested event set) will unblock port_getn right away
// because of the I/O readiness notification.
pd.user = 0
tp := taggedPointerPack(unsafe.Pointer(pd), pd.fdseq.Load())
// Note that this won't work on a 32-bit system,
// as taggedPointer is always 64-bits but uintptr will be 32 bits.
// Fortunately we only support Solaris on amd64.
if goarch.PtrSize != 8 {
throw("runtime: netpollopen: unsupported pointer size")
}
r := port_associate(portfd, _PORT_SOURCE_FD, fd, 0, uintptr(tp))
unlock(&pd.lock)
return r
}
func netpollclose(fd uintptr) int32 {
return port_dissociate(portfd, _PORT_SOURCE_FD, fd)
}
// Updates the association with a new set of interested events. After
// this call, port_getn will return one and only one event for that
// particular descriptor, so this function needs to be called again.
func netpollupdate(pd *pollDesc, set, clear uint32) {
if pd.info().closing() {
return
}
old := pd.user
events := (old & ^clear) | set
if old == events {
return
}
tp := taggedPointerPack(unsafe.Pointer(pd), pd.fdseq.Load())
if events != 0 && port_associate(portfd, _PORT_SOURCE_FD, pd.fd, events, uintptr(tp)) != 0 {
print("runtime: port_associate failed (errno=", errno(), ")\n")
throw("runtime: netpollupdate failed")
}
pd.user = events
}
// subscribe the fd to the port such that port_getn will return one event.
func netpollarm(pd *pollDesc, mode int) {
lock(&pd.lock)
switch mode {
case 'r':
netpollupdate(pd, _POLLIN, 0)
case 'w':
netpollupdate(pd, _POLLOUT, 0)
default:
throw("runtime: bad mode")
}
unlock(&pd.lock)
}
// netpollBreak interrupts a port_getn wait.
func netpollBreak() {
// Failing to cas indicates there is an in-flight wakeup, so we're done here.
if !netpollWakeSig.CompareAndSwap(0, 1) {
return
}
// Use port_alert to put portfd into alert mode.
// This will wake up all threads sleeping in port_getn on portfd,
// and cause their calls to port_getn to return immediately.
// Further, until portfd is taken out of alert mode,
// all calls to port_getn will return immediately.
if port_alert(portfd, _PORT_ALERT_UPDATE, _POLLHUP, uintptr(unsafe.Pointer(&portfd))) < 0 {
if e := errno(); e != _EBUSY {
println("runtime: port_alert failed with", e)
throw("runtime: netpoll: port_alert failed")
}
}
}
// netpoll checks for ready network connections.
// Returns a list of goroutines that become runnable,
// and a delta to add to netpollWaiters.
// This must never return an empty list with a non-zero delta.
//
// delay < 0: blocks indefinitely
// delay == 0: does not block, just polls
// delay > 0: block for up to that many nanoseconds
func netpoll(delay int64) (gList, int32) {
if portfd == -1 {
return gList{}, 0
}
var wait *timespec
var ts timespec
if delay < 0 {
wait = nil
} else if delay == 0 {
wait = &ts
} else {
ts.setNsec(delay)
if ts.tv_sec > 1e6 {
// An arbitrary cap on how long to wait for a timer.
// 1e6 s == ~11.5 days.
ts.tv_sec = 1e6
}
wait = &ts
}
var events [128]portevent
retry:
var n uint32 = 1
r := port_getn(portfd, &events[0], uint32(len(events)), &n, wait)
e := errno()
if r < 0 && e == _ETIME && n > 0 {
// As per port_getn(3C), an ETIME failure does not preclude the
// delivery of some number of events. Treat a timeout failure
// with delivered events as a success.
r = 0
}
if r < 0 {
if e != _EINTR && e != _ETIME {
print("runtime: port_getn on fd ", portfd, " failed (errno=", e, ")\n")
throw("runtime: netpoll failed")
}
// If a timed sleep was interrupted and there are no events,
// just return to recalculate how long we should sleep now.
if delay > 0 {
return gList{}, 0
}
goto retry
}
var toRun gList
delta := int32(0)
for i := 0; i < int(n); i++ {
ev := &events[i]
if ev.portev_source == _PORT_SOURCE_ALERT {
if ev.portev_events != _POLLHUP || unsafe.Pointer(ev.portev_user) != unsafe.Pointer(&portfd) {
throw("runtime: netpoll: bad port_alert wakeup")
}
if delay != 0 {
// Now that a blocking call to netpoll
// has seen the alert, take portfd
// back out of alert mode.
// See the comment in netpollBreak.
if port_alert(portfd, 0, 0, 0) < 0 {
e := errno()
println("runtime: port_alert failed with", e)
throw("runtime: netpoll: port_alert failed")
}
netpollWakeSig.Store(0)
}
continue
}
if ev.portev_events == 0 {
continue
}
tp := taggedPointer(uintptr(unsafe.Pointer(ev.portev_user)))
pd := (*pollDesc)(tp.pointer())
if pd.fdseq.Load() != tp.tag() {
continue
}
var mode, clear int32
if (ev.portev_events & (_POLLIN | _POLLHUP | _POLLERR)) != 0 {
mode += 'r'
clear |= _POLLIN
}
if (ev.portev_events & (_POLLOUT | _POLLHUP | _POLLERR)) != 0 {
mode += 'w'
clear |= _POLLOUT
}
// To effect edge-triggered events, we need to be sure to
// update our association with whatever events were not
// set with the event. For example if we are registered
// for POLLIN|POLLOUT, and we get POLLIN, besides waking
// the goroutine interested in POLLIN we have to not forget
// about the one interested in POLLOUT.
if clear != 0 {
lock(&pd.lock)Bash sed -n '320,360p' netpoll_solaris.go
sed -n '320,360p' netpoll_solaris.go
Rest of netpoll solaris
lock(&pd.lock)
netpollupdate(pd, 0, uint32(clear))
unlock(&pd.lock)
}
if mode != 0 {
// TODO(mikio): Consider implementing event
// scanning error reporting once we are sure
// about the event port on SmartOS.
//
// See golang.org/x/issue/30840.
delta += netpollready(&toRun, pd, mode)
}
}
return toRun, delta
}Bash cd /tmp; q() { curl -s -G "https://api.github.com/search/issues" --data-urlencode "q=$1" | python3 -c " ...
cd /tmp; q() { curl -s -G "https://api.github.com/search/issues" --data-urlencode "q=$1" | python3 -c "
import json,sys
d=json.load(sys.stdin)
print('total',d.get('total_count'))
for i in d.get('items',[])[:15]:
print(i['number'], i['state'], i['title'])
"; }
echo "--- A"; q 'repo:golang/go solaris netpoll in:title'
echo "--- B"; q 'repo:golang/go "blocked in Close" solaris in:body'
echo "--- C"; q 'repo:golang/go solaris hang write netpoll'Search for related solaris issues
--- A total 1 7410 closed net: fix spurious netpoll "failed to associate" error on Solaris --- B total 0 --- C total 6 34988 open syscall: memory corruption when forking on OpenBSD, NetBSD, AIX, and Solaris 53289 open runtime: frequent GC related asserts seen on illumos based OSes (but only on AMD!) 58901 open all: test failures on `linux/s390x` 57999 open all: test failures on `netbsd` blocked on waiting for subprocesses 61666 closed x/build: add LUCI solaris-amd64 builder 48751 closed x/tools/gopls: VS code hangs when working on a project that uses bazel. Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /home/claude/golang_go_78438/go/src/net/http && sed -n '1203,1295p' transport_test.go | cat -n | awk '{printf "%d\t%s...
cd /home/claude/golang_go_78438/go/src/net/http && sed -n '1203,1295p' transport_test.go | cat -n | awk '{printf "%d\t%s\n", $1+1202, substr($0, index($0,$2))}' | sed -n '1,95p'Line-numbered test source
1203 func TestTransportGzip(t *testing.T) { run(t, testTransportGzip, http3SkippedMode) }
1204 func testTransportGzip(t *testing.T, mode testMode) {
1205 const testString = "The test string aaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaa"
1206 const nRandBytes = 1024 * 1024
1207 ts := newClientServerTest(t, mode, HandlerFunc(func(rw ResponseWriter, req *Request) {
1208 if req.Method == "HEAD" {
1209 if g := req.Header.Get("Accept-Encoding"); g != "" {
1210 t.Errorf("HEAD request sent with Accept-Encoding of %q; want none", g)
1211 }
1212 return
1213 }
1214 if g, e := req.Header.Get("Accept-Encoding"), "gzip"; g != e {
1215 t.Errorf("Accept-Encoding = %q, want %q", g, e)
1216 }
1217 rw.Header().Set("Content-Encoding", "gzip")
1218 16
1219 var w io.Writer = rw
1220 var buf bytes.Buffer
1221 if req.FormValue("chunked") == "0" {
1222 w = &buf
1223 defer io.Copy(rw, &buf)
1224 defer func() {
1225 rw.Header().Set("Content-Length", strconv.Itoa(buf.Len()))
1226 }()
1227 }
1228 gz := gzip.NewWriter(w)
1229 gz.Write([]byte(testString))
1230 if req.FormValue("body") == "large" {
1231 io.CopyN(gz, rand.Reader, nRandBytes)
1232 }
1233 gz.Close()
1234 })).ts
1235 c := ts.Client()
1236 34
1237 for _, chunked := range []string{"1", "0"} {
1238 // First fetch something large, but only read some of it.
1239 res, err := c.Get(ts.URL + "/?body=large&chunked=" + chunked)
1240 if err != nil {
1241 t.Fatalf("large get: %v", err)
1242 }
1243 buf := make([]byte, len(testString))
1244 n, err := io.ReadFull(res.Body, buf)
1245 if err != nil {
1246 t.Fatalf("partial read of large response: size=%d, %v", n, err)
1247 }
1248 if e, g := testString, string(buf); e != g {
1249 t.Errorf("partial read got %q, expected %q", g, e)
1250 }
1251 res.Body.Close()
1252 // Read on the body, even though it's closed
1253 n, err = res.Body.Read(buf)
1254 if n != 0 || err == nil {
1255 t.Errorf("expected error post-closed large Read; got = %d, %v", n, err)
1256 }
1257 55
1258 // Then something small.
1259 res, err = c.Get(ts.URL + "/?chunked=" + chunked)
1260 if err != nil {
1261 t.Fatal(err)
1262 }
1263 body, err := io.ReadAll(res.Body)
1264 if err != nil {
1265 t.Fatal(err)
1266 }
1267 if g, e := string(body), testString; g != e {
1268 t.Fatalf("body = %q; want %q", g, e)
1269 }
1270 if g, e := res.Header.Get("Content-Encoding"), ""; g != e {
1271 t.Fatalf("Content-Encoding = %q; want %q", g, e)
1272 }
1273 71
1274 // Read on the body after it's been fully read:
1275 n, err = res.Body.Read(buf)
1276 if n != 0 || err == nil {
1277 t.Errorf("expected Read error after exhausted reads; got %d, %v", n, err)
1278 }
1279 res.Body.Close()
1280 n, err = res.Body.Read(buf)
1281 if n != 0 || err == nil {
1282 t.Errorf("expected Read error after Close; got %d, %v", n, err)
1283 }
1284 }
1285 83
1286 // And a HEAD request too, because they're always weird.
1287 res, err := c.Head(ts.URL)
1288 if err != nil {
1289 t.Fatalf("Head: %v", err)
1290 }
1291 if res.StatusCode != 200 {
1292 t.Errorf("Head status=%d; want=200", res.StatusCode)
1293 }
1294 }
1295 93Bash cd /home/claude/golang_go_78438/go && git log --oneline -12 --since=2025-06-01 -- src/runtime/netpoll_solaris.go src/run...
cd /home/claude/golang_go_78438/go && git log --oneline -12 --since=2025-06-01 -- src/runtime/netpoll_solaris.go src/runtime/netpoll.go src/internal/poll/fd_unix.go src/internal/poll/fd_poll_runtime.go 2>&1; echo "=== httptest"; git log --oneline -8 -- src/net/http/httptest/server.go; echo "=== test"; git log --oneline -5 -- src/net/http/transport_test.go | head
Recent commits to netpoll/httptest
25de5ebdcd internal/poll: document why zero-byte reads do not wait for readability 49c1da474d internal/itoa, internal/runtime/strconv: delete === httptest ca4a27f0b1 net/http: rework HTTP/3 registration to allow using a fake network 813b317cc9 net/http/httptest: add NewTestServer with in-memory network 2f3c778b23 net/http: add support for running HTTP tests against HTTP/3 f2db0dca0b net/http/httptest: redirect example.com requests to server 69c74c9ea9 net/http/httptest: add comment to Server.Client() about Server.URL 1d45a7ef56 net: add available godoc link f5ec2e46d9 net/http/httptest: remove unnecessary creation of http.Transport 0dfb22ed70 all: use ^TestName$ regular pattern for invoking a single test === test 5cfba704d0 net/http/internal/http2: use nettest for fake networking 857fa809d2 net/http: enable HTTP/2 on http2.ConfigureServer 5005a2068b net/http: prevent blocking when reading and closing gzipReader concurrently c22f92a751 net/http: fix hang in TestTransportClosesBodyOnError/h3 65d5c5f6dd net/http: resolve data race in TestTransportReadToEndReusesConn
Bash git log -1 --format='%H %ad %s%n%n%b' --date=short 5005a2068b | head -40
git log -1 --format='%H %ad %s%n%n%b' --date=short 5005a2068b | head -40
Show gzipReader commit
5005a2068bc525f0d209cea43601a7d3f0a8cde5 2026-06-08 net/http: prevent blocking when reading and closing gzipReader concurrently This backports the fix for an issue found while adding gzip support for HTTP/3 in CL 785120. Change-Id: Idf681f595c9d471db76b5ce3879a83a16a6a6964 Reviewed-on: https://go-review.googlesource.com/c/go/+/788480 Reviewed-by: Cherry Mui <cherryyz@google.com> Reviewed-by: Damien Neil <dneil@google.com> LUCI-TryBot-Result: golang-scoped@luci-project-accounts.iam.gserviceaccount.com <golang-scoped@luci-project-accounts.iam.gserviceaccount.com>
Bash cd /home/claude/golang_go_78438/go/src/net/http/httptest && sed -n '495,600p' server.go
cd /home/claude/golang_go_78438/go/src/net/http/httptest && sed -n '495,600p' server.go
Read httptest Server.Close
// Force-close any idle connections (those between
// requests) and new connections (those which connected
// but never sent a request). StateNew connections are
// super rare and have only been seen (in
// previously-flaky tests) in the case of
// socket-late-binding races from the http Client
// dialing this server and then getting an idle
// connection before the dial completed. There is thus
// a connected connection in StateNew with no
// associated Request. We only close StateIdle and
// StateNew because they're not doing anything. It's
// possible StateNew is about to do something in a few
// milliseconds, but a previous CL to check again in a
// few milliseconds wasn't liked (early versions of
// https://golang.org/cl/15151) so now we just
// forcefully close StateNew. The docs for Server.Close say
// we wait for "outstanding requests", so we don't close things
// in StateActive.
if st == http.StateIdle || st == http.StateNew {
s.closeConn(c)
}
}
// If this server doesn't shut down in 5 seconds, tell the user why.
t := time.AfterFunc(5*time.Second, s.logCloseHangDebugInfo)
defer t.Stop()
}
s.mu.Unlock()
// Not part of httptest.Server's correctness, but assume most
// users of httptest.Server will be using the standard
// transport, so help them out and close any idle connections for them.
if t, ok := http.DefaultTransport.(closeIdleTransport); ok {
t.CloseIdleConnections()
}
// Also close the client idle connections.
if s.client != nil {
if t, ok := s.client.Transport.(closeIdleTransport); ok {
t.CloseIdleConnections()
}
}
s.wg.Wait()
}
func (s *Server) logCloseHangDebugInfo() {
s.mu.Lock()
defer s.mu.Unlock()
var buf strings.Builder
buf.WriteString("httptest.Server blocked in Close after 5 seconds, waiting for connections:\n")
for c, st := range s.conns {
fmt.Fprintf(&buf, " %T %p %v in state %v\n", c, c, c.RemoteAddr(), st)
}
log.Print(buf.String())
}
// CloseClientConnections closes any open HTTP connections to the test Server.
func (s *Server) CloseClientConnections() {
s.mu.Lock()
nconn := len(s.conns)
ch := make(chan struct{}, nconn)
for c := range s.conns {
go s.closeConnChan(c, ch)
}
s.mu.Unlock()
// Wait for outstanding closes to finish.
//
// Out of paranoia for making a late change in Go 1.6, we
// bound how long this can wait, since golang.org/issue/14291
// isn't fully understood yet. At least this should only be used
// in tests.
timer := time.NewTimer(5 * time.Second)
defer timer.Stop()
for i := 0; i < nconn; i++ {
select {
case <-ch:
case <-timer.C:
// Too slow. Give up.
return
}
}
}
// Certificate returns the certificate used by the server, or nil if
// the server doesn't use TLS.
func (s *Server) Certificate() *x509.Certificate {
return s.certificate
}
// Client returns an HTTP client configured for making requests to the server.
// It is configured to trust the server's TLS test certificate and will
// close its idle connections on [Server.Close].
func (s *Server) Client() *http.Client {
if s.t != nil {
s.startOnce.Do(s.startFakeNet)
}
return s.client
}
func (s *Server) goServe(li net.Listener) {
s.wg.Add(1)
go func() {
defer s.wg.Done()
s.Config.Serve(li)
}()
}Bash cd /tmp; q() { curl -s -G "https://api.github.com/search/issues" --data-urlencode "q=$1" | python3 -c " ...
cd /tmp; q() { curl -s -G "https://api.github.com/search/issues" --data-urlencode "q=$1" | python3 -c "
import json,sys
d=json.load(sys.stdin)
print('total',d.get('total_count'))
for i in d.get('items',[])[:12]:
print(i['number'], i['state'], i['title'])
"; }
echo "--- fusion"; q 'repo:golang/go solaris fusion tcp'
echo "--- solaris write block"; q 'repo:golang/go repo:golang/go solaris "test timed out" net/http'
echo "--- solaris flakes"; q 'repo:golang/go label:Automation solaris in:title'
echo "--- gzip test"; q 'repo:golang/go TestTransportGzip in:title'Search related issues
--- fusion total 0 --- solaris write block total 21 60152 open all: test timeouts on the `oraclerel` builder 62357 open net/http: apparent deadlock in TestTransportDecrementConnWhenIdleConnRemoved 78647 open net/http: TestClientRedirectTypes/h3 failures 65311 open x/build: timeout on LUCI openbsd-amd64 builder 56033 open x/net/nettest: TestTestConn failures 57999 open all: test failures on `netbsd` blocked on waiting for subprocesses 55162 open all: test failures using proxy.golang.org 58901 open all: test failures on `linux/s390x` 66474 open x/build: build and test failures on openbsd-amd64 builders 74187 open all: umbrella for flakes on openbsd/ppc64 55167 open runtime: corruption in tests 55308 closed runtime: unrecognized failures --- solaris flakes total 0 --- gzip test total 1 78438 open net/http: TestTransportGzip/h1 failures Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; curl -sL https://api.github.com/repos/golang/go/issues/60152 | python3 -c " ...
cd /tmp; curl -sL https://api.github.com/repos/golang/go/issues/60152 | python3 -c " import json,sys d=json.load(sys.stdin) print(d['title'], d['state'], [l['name'] for l in d['labels']]) print(d['body'][:4000]) "
Read issue 60152
all: test timeouts on the `oraclerel` builder open ['OS-Solaris', 'Builders', 'NeedsInvestigation'] ``` #!watchflakes default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT") ``` (CC @golang/solaris @rorth) Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; curl -sL "https://api.github.com/repos/golang/go/issues/60152/comments?per_page=100&page=1" -o c60152.json; pyt...
cd /tmp; curl -sL "https://api.github.com/repos/golang/go/issues/60152/comments?per_page=100&page=1" -o c60152.json; python3 -c "
import json
d=json.load(open('c60152.json'))
print(len(d))
for c in d[-12:]:
print('===', c['user']['login'], c['created_at'])
print(c['body'][:1500])
"Read recent comments on 60152
56
=== gopherbot 2025-10-08T20:43:23Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-10-07 19:46 gotip-solaris-amd64 go@f6f4e8b3 cmd/compile.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8701551815601747697">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-10-08T20:02:49Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics626321199/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics626321199/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics626321199/001/tmp
# Test of the linker's script test harness. (20.490s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
-
=== gopherbot 2025-10-09T04:42:41Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-10-08 20:44 gotip-solaris-amd64 go@d4830c61 cmd/link.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8701521331565140321">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-10-09T04:02:09Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2016912786/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2016912786/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2016912786/001/tmp
# Test of the linker's script test harness. (3.625s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
--
=== gopherbot 2025-10-13T21:43:29Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-10-13 19:19 gotip-solaris-amd64 go@9b8742f2 cmd/compile.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8701098121484085265">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-10-13T19:54:28Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2759434483/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2759434483/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2759434483/001/tmp
# Test of the linker's script test harness. (18.988s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
=== gopherbot 2025-10-21T20:02:11Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-10-21 17:45 gotip-solaris-amd64 go@8a6c64f4 cmd/link.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8700377646905006817">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-10-21T18:52:46Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics4256542120/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics4256542120/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics4256542120/001/tmp
# Test of the linker's script test harness. (2.580s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
--
=== gopherbot 2025-10-22T16:39:03Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-10-22 15:20 gotip-solaris-amd64 go@4bdb55b5 cmd/compile.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8700297819931751409">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-10-22T15:29:40Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics802955447/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics802955447/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics802955447/001/tmp
# Test of the linker's script test harness. (24.174s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
-
=== gopherbot 2025-10-25T12:14:51Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-10-24 14:23 gotip-solaris-amd64 go@c445a61e cmd/compile.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8700041004301771649">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-10-25T11:54:12Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics1583130119/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics1583130119/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics1583130119/001/tmp
# Test of the linker's script test harness. (18.613s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
=== gopherbot 2025-10-28T11:55:53Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-10-28 02:06 gotip-solaris-amd64 go@53be7863 cmd/link.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8699770826930384673">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-10-28T11:33:39Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics4021963500/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics4021963500/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics4021963500/001/tmp
# Test of the linker's script test harness. (1.718s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
--
=== gopherbot 2025-11-13T02:00:21Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-11-12 21:54 gotip-solaris-amd64 go@f03d06ec cmd/link.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8698369084841711873">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-11-12T23:09:07Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2024084071/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2024084071/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics2024084071/001/tmp
# Test of the linker's script test harness. (2.368s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
--
=== gopherbot 2025-11-22T23:55:15Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-11-21 18:51 gotip-solaris-amd64 go@c851827c cmd/compile.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8697460740114348129">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:223: 2025-11-22T23:30:53Z
run.go:225: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics3946161009/001
run.go:232:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics3946161009/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics3946161009/001/tmp
# Test of the linker's script test harness. (26.430s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:232: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
=== gopherbot 2025-12-05T16:38:04Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-12-05 01:35 gotip-solaris-amd64 go@91267f0a cmd/compile.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8696363437122106881">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:226: 2025-12-05T02:04:15Z
run.go:228: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics3821665916/001
run.go:235:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics3821665916/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics3821665916/001/tmp
# Test of the linker's script test harness. (13.615s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:235: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
=== gopherbot 2025-12-18T18:23:59Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2025-12-11 16:34 gotip-solaris-amd64 go@e38c38f0 cmd/link.TestScript/script_test_basics (<a href="https://ci.chromium.org/b/8695727132333360593">log</a>)</summary>
=== RUN TestScript/script_test_basics
=== PAUSE TestScript/script_test_basics
=== CONT TestScript/script_test_basics
run.go:226: 2025-12-12T02:44:49Z
run.go:228: $WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics1916971993/001
run.go:235:
BOTO_CONFIG=/opt/golang/swarm/.swarming/w/ir/x/a/gsutil-bbagent/.boto
CIPD_ARCHITECTURE=amd64
CIPD_CACHE_DIR=/opt/golang/swarm/.swarming/w/ir/cache/cipd_cache
CIPD_PROTOCOL=v2
...
WORK=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics1916971993/001
TMPDIR=/opt/golang/swarm/.swarming/w/ir/x/t/TestScriptscript_test_basics1916971993/001/tmp
# Test of the linker's script test harness. (2.991s)
> go build
> [!cgo] skip
[condition not met]
> cc -c testdata/mumble.c
run.go:235: FAIL: testdata/script/script_test_basics.txt:6: cc -c testdata/mumble.c: exec: WaitDelay expired before I/O complete
--
=== gopherbot 2026-05-22T05:50:32Z
Found new dashboard test flakes for:
#!watchflakes
default <- goos == "solaris" && goarch == "amd64" && (`test timed out` || `failed to start: context deadline exceeded` || `WaitDelay expired before I/O complete` || status == "ABORT")
<details><summary>2026-05-22 02:39 x_crypto-go1.26-solaris-amd64 crypto@d2fe592d release-branch.go1.26@26c2ac5f x/crypto/ssh.TestMuxChannelOverflow [SKIP] (<a href="https://ci.chromium.org/b/8681133594973173889">log</a>)</summary>
=== RUN TestMuxChannelOverflow
panic: test timed out after 10m0s
running tests:
TestMuxChannelOverflow (9m54s)
goroutine 1693 [running]:
testing.(*M).startAlarm.func1()
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/testing/testing.go:2802 +0x354
created by time.goFunc
/opt/golang/swarm/.swarming/w/ir/x/w/goroot/src/time/sleep.go:215 +0x2d
...
created by golang.org/x/crypto/ssh.newServerTransport in goroutine 1359
/opt/golang/swarm/.swarming/w/ir/x/w/targetrepo2992368679/ssh/handshake.go:180 +0x165
goroutine 1467 [chan receive, 9 minutes]:
golang.org/x/crypto/ssh.(*handshakeTransport).readOnePacket(0x1e672a9cc680, 0x1)
/opt/golang/swarm/.swarming/w/ir/x/w/targetrepo2992368679/ssh/handshake.go:458 +0x22a
golang.org/x/crypto/ssh.TestStrictKEXMixed.func2()
/opt/golang/swarm/.swarming/w/ir/x/w/targetrepo2992368679/ssh/handshake_test.go:1194 +0x2b
created by golang.org/x/crypto/ssh.TestStrictKEXMixed in goroutine 1464
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; for id in 8686112892523717025 8681022390813019185 8675856801014264801 8676093236579846177; do ...
cd /tmp; for id in 8686112892523717025 8681022390813019185 8675856801014264801 8676093236579846177; do
( curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/$id/+/u/step/11/log/3?format=raw" | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('s_$id.txt','w').write(''.join(out))
" ) &
done; wait; ls -l s_*.txtFetch remaining solaris logs in parallel
[1] Done ( curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/$id/+/u/step/11/log/3?format=raw" | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('s_$id.txt','w').write(''.join(out))
" )
[3]- Done ( curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/$id/+/u/step/11/log/3?format=raw" | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('s_$id.txt','w').write(''.join(out))
" )
[2]- Done ( curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/$id/+/u/step/11/log/3?format=raw" | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('s_$id.txt','w').write(''.join(out))
" )
[4]+ Done ( curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/$id/+/u/step/11/log/3?format=raw" | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('s_$id.txt','w').write(''.join(out))
" )
-rw-rw-r-- 1 claude claude 273143 Aug 25 08:22 s_8675856801014264801.txt
-rw-rw-r-- 1 claude claude 272323 Aug 25 08:22 s_8676093236579846177.txt
-rw-rw-r-- 1 claude claude 265195 Aug 25 08:22 s_8681022390813019185.txt
-rw-rw-r-- 1 claude claude 266625 Aug 25 08:22 s_8686112892523717025.txt
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; for f in s_*.txt; do echo "=========== $f"; grep -n "blocked in Close" $f | head -3; grep -n "running tests:" -...
cd /tmp; for f in s_*.txt; do echo "=========== $f"; grep -n "blocked in Close" $f | head -3; grep -n "running tests:" -A3 $f | head -6; python3 - "$f" <<'EOF'
import re,sys
s=open(sys.argv[1]).read()
i=s.index('panic: test timed out')
for b in re.split(r'\n\n', s[i:]):
if 'net/http_test.testTransportGzip' in b or 'FD).Write' in b or 'httptest.(*Server).Close' in b:
lines=[l for l in b.strip().split('\n') if not l.startswith('\t')]
print(' ', lines[0])
print(' ', ' <- '.join(l.split('(')[0] for l in lines[1:9]))
EOF
doneCompare signatures across solaris logs
=========== s_8675856801014264801.txt
5211:2026/07/19 12:33:36 httptest.Server blocked in Close after 5 seconds, waiting for connections:
5214: running tests:
5215- TestTransportGzip/h1 (9m57s)
5216-
5217-goroutine 18452 gp=0x7675f65f860 m=11 mp=0x7675e82d008 [running]:
goroutine 4335 gp=0x7675e8dd680 m=nil [sync.WaitGroup.Wait, 9 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.semacquire1 <- sync.runtime_SemacquireWaitGroup <- sync. <- net/http/httptest. <- net/http_test. <- net/http_test.newClientServerTest.func5
goroutine 15777 gp=0x7675ea12f00 m=nil [IO wait, 9 minutes]:
runtime.gopark <- runtime.netpollblock <- internal/poll.runtime_pollWait <- internal/poll. <- internal/poll. <- internal/poll. <- net. <- net.
=========== s_8676093236579846177.txt
5188:2026/07/16 23:03:27 httptest.Server blocked in Close after 5 seconds, waiting for connections:
5191: running tests:
5192- TestTransportGzip/h1 (9m58s)
5193-
5194-goroutine 18396 gp=0x212cb55e2f00 m=9 mp=0x212cb5080008 [running]:
goroutine 3330 gp=0x212cb5232960 m=nil [sync.WaitGroup.Wait, 9 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.semacquire1 <- sync.runtime_SemacquireWaitGroup <- sync. <- net/http/httptest. <- net/http_test. <- net/http_test.newClientServerTest.func5
goroutine 6680 gp=0x212cb53890e0 m=nil [IO wait, 9 minutes]:
runtime.gopark <- runtime.netpollblock <- internal/poll.runtime_pollWait <- internal/poll. <- internal/poll. <- internal/poll. <- net. <- net.
=========== s_8681022390813019185.txt
5091:2026/05/23 12:07:20 httptest.Server blocked in Close after 5 seconds, waiting for connections:
5094: running tests:
5095- TestTransportGzip/h1 (8m58s)
5096-
5097-goroutine 18375 gp=0xa5cc074d20 m=0 mp=0xfb4bc0 [running]:
goroutine 7437 gp=0xa5cc6065a0 m=nil [IO wait, 8 minutes]:
runtime.gopark <- runtime.netpollblock <- internal/poll.runtime_pollWait <- internal/poll. <- internal/poll. <- internal/poll. <- net. <- net.
goroutine 3309 gp=0xa5cc1332c0 m=nil [sync.WaitGroup.Wait, 8 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.semacquire1 <- sync.runtime_SemacquireWaitGroup <- sync. <- net/http/httptest. <- net/http_test. <- net/http_test.newClientServerTest.func5
=========== s_8686112892523717025.txt
5101:2026/03/28 06:57:06 httptest.Server blocked in Close after 5 seconds, waiting for connections:
5105: running tests:
5106- TestTransportGzip/h1 (8m58s)
5107-
5108-goroutine 18555 gp=0x28015820b4a0 m=21 mp=0x28015827d008 [running]:
goroutine 8341 gp=0x2801584f21e0 m=nil [IO wait, 8 minutes]:
runtime.gopark <- runtime.netpollblock <- internal/poll.runtime_pollWait <- internal/poll. <- internal/poll. <- internal/poll. <- net. <- net.
goroutine 3445 gp=0x280157ed83c0 m=nil [sync.WaitGroup.Wait, 8 minutes]:
runtime.gopark <- runtime.goparkunlock <- runtime.semacquire1 <- sync.runtime_SemacquireWaitGroup <- sync. <- net/http/httptest. <- net/http_test. <- net/http_test.newClientServerTest.func2
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; for f in s_*.txt; do echo "== $f"; sed -n "$(grep -n 'blocked in Close' $f | head -1 | cut -d: -f1),+3p" $f; py...
cd /tmp; for f in s_*.txt; do echo "== $f"; sed -n "$(grep -n 'blocked in Close' $f | head -1 | cut -d: -f1),+3p" $f; python3 - "$f" <<'EOF'
import re,sys
s=open(sys.argv[1]).read(); i=s.index('panic: test timed out')
for b in re.split(r'\n\n', s[i:]):
if 'FD).Write' in b:
for l in b.split('\n'):
if 'FD).Write' in l or 'chunkWriter).Write' in l or 'testTransportGzip' in l or 'deferwrap' in l:
print(' ', l.strip())
EOF
doneDetails of blocked write in each solaris log
== s_8675856801014264801.txt
2026/07/19 12:33:36 httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0x7675ecaa8c8 127.0.0.1:49964 in state active
panic: test timed out after 10m0s
running tests:
internal/poll.(*FD).Write(0x767602cef00, {0x76762060f97, 0xff10d, 0x11f069})
net.(*netFD).Write(0x767602cef00, {0x76762060f97?, 0x0?, 0x0?})
net/http.(*chunkWriter).Write(0x76760cb75d8, {0x76762060000, 0x1000a4, 0x120000})
net/http_test.testTransportGzip.func1.deferwrap1()
net/http_test.testTransportGzip.func1({0xf67d98, 0x76760cb7590}, 0x767610ac500)
== s_8676093236579846177.txt
2026/07/16 23:03:27 httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0x212cb50222f8 127.0.0.1:60535 in state active
panic: test timed out after 10m0s
running tests:
internal/poll.(*FD).Write(0x212cb5b17800, {0x212cb6464f97, 0xff10d, 0x11f069})
net.(*netFD).Write(0x212cb5b17800, {0x212cb6464f97?, 0x0?, 0x0?})
net/http.(*chunkWriter).Write(0x212cb5914220, {0x212cb6464000, 0x1000a4, 0x120000})
net/http_test.testTransportGzip.func1.deferwrap1()
net/http_test.testTransportGzip.func1({0xf64848, 0x212cb59141e0}, 0x212cb5904000)
== s_8681022390813019185.txt
2026/05/23 12:07:20 httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0xa5ccc8e198 127.0.0.1:53482 in state active
panic: test timed out after 9m0s
running tests:
internal/poll.(*FD).Write(0xa5cca65800, {0xa5cdf6cf97, 0xff10d, 0x11f069})
net.(*netFD).Write(0xa5cca65800, {0xa5cdf6cf97?, 0x0?, 0x0?})
net/http.(*chunkWriter).Write(0xa5cdbba6d0, {0xa5cdf6c000, 0x1000a4, 0x120000})
net/http_test.testTransportGzip.func1.deferwrap1()
net/http_test.testTransportGzip.func1({0xf31f98, 0xa5cdbba690}, 0xa5cc262640)
== s_8686112892523717025.txt
2026/03/28 06:57:06 httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0x280157d88320 127.0.0.1:59305 in state active
--- PASS: TestOmitHTTP2Vet (7.97s)
panic: test timed out after 9m0s
internal/poll.(*FD).Write(0x28015b781400, {0x28015c100f97, 0xff1f7, 0x1ff069})
net.(*netFD).Write(0x28015b781400, {0x28015c100f97?, 0x0?, 0x0?})
net/http.(*chunkWriter).Write(0x2801592153f0, {0x28015c100000, 0x10018e, 0x200000})
net/http_test.testTransportGzip.func1.deferwrap1()
net/http_test.testTransportGzip.func1({0xdfdbf0, 0x2801592153b0}, 0x28015bbafcc0)
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /home/claude/golang_go_78438/go/src/net/http && grep -n "startBackgroundRead\|abortPendingRead\|func (cr \*connReader...
cd /home/claude/golang_go_78438/go/src/net/http && grep -n "startBackgroundRead\|abortPendingRead\|func (cr \*connReader) backgroundRead" server.go
Find background read call sites
322: c.r.abortPendingRead()
731:func (cr *connReader) startBackgroundRead() {
745:func (cr *connReader) backgroundRead() {
775: // another goroutine calling abortPendingRead.
785:func (cr *connReader) abortPendingRead() {
1706: w.conn.r.abortPendingRead()
1944: inFlightResponse.conn.r.abortPendingRead()
2120: w.reqBody.registerOnHitEOF(w.conn.r.startBackgroundRead)
2122: w.conn.r.startBackgroundRead()Bash sed -n '310,330p;725,800p' server.go; echo "=== 1700"; sed -n '1698,1712p' server.go
sed -n '310,330p;725,800p' server.go; echo "=== 1700"; sed -n '1698,1712p' server.go
Read connReader background read code
func (c *conn) hijacked() bool {
c.mu.Lock()
defer c.mu.Unlock()
return c.hijackedv
}
// c.mu must be held.
func (c *conn) hijackLocked() (rwc net.Conn, buf *bufio.ReadWriter, err error) {
if c.hijackedv {
return nil, nil, ErrHijacked
}
c.r.abortPendingRead()
c.hijackedv = true
rwc = c.rwc
rwc.SetDeadline(time.Time{})
if c.r.hasByte {
if _, err := c.bufr.Peek(c.bufr.Buffered() + 1); err != nil {
return nil, nil, fmt.Errorf("unexpected Peek failure reading buffered byte: %v", err)
func (cr *connReader) releaseConn() {
cr.lock()
defer cr.unlock()
cr.conn = nil
}
func (cr *connReader) startBackgroundRead() {
cr.lock()
defer cr.unlock()
if cr.inRead {
panic("invalid concurrent Body.Read call")
}
if cr.hasByte {
return
}
cr.inRead = true
cr.rwc.SetReadDeadline(time.Time{})
go cr.backgroundRead()
}
func (cr *connReader) backgroundRead() {
n, err := cr.rwc.Read(cr.byteBuf[:])
cr.lock()
if n == 1 {
cr.hasByte = true
// We were past the end of the previous request's body already
// (since we wouldn't be in a background read otherwise), so
// this is a pipelined HTTP request. Prior to Go 1.11 we used to
// send on the CloseNotify channel and cancel the context here,
// but the behavior was documented as only "may", and we only
// did that because that's how CloseNotify accidentally behaved
// in very early Go releases prior to context support. Once we
// added context support, people used a Handler's
// Request.Context() and passed it along. Having that context
// cancel on pipelined HTTP requests caused problems.
// Fortunately, almost nothing uses HTTP/1.x pipelining.
// Unfortunately, apt-get does, or sometimes does.
// New Go 1.11 behavior: don't fire CloseNotify or cancel
// contexts on pipelined requests. Shouldn't affect people, but
// fixes cases like Issue 23921. This does mean that a client
// closing their TCP connection after sending a pipelined
// request won't cancel the context, but we'll catch that on any
// write failure (in checkConnErrorWriter.Write).
// If the server never writes, yes, there are still contrived
// server & client behaviors where this fails to ever cancel the
// context, but that's kinda why HTTP/1.x pipelining died
// anyway.
}
if ne, ok := err.(net.Error); ok && cr.aborted && ne.Timeout() {
// Ignore this error. It's the expected error from
// another goroutine calling abortPendingRead.
} else if err != nil {
cr.handleReadErrorLocked(err)
}
cr.aborted = false
cr.inRead = false
cr.unlock()
cr.cond.Broadcast()
}
func (cr *connReader) abortPendingRead() {
cr.lock()
defer cr.unlock()
if !cr.inRead {
return
}
cr.aborted = true
cr.rwc.SetReadDeadline(aLongTimeAgo)
for cr.inRead {
cr.cond.Wait()
}
cr.rwc.SetReadDeadline(time.Time{})
}
func (cr *connReader) setReadLimit(remain int64) { cr.remain = remain }
func (cr *connReader) setInfiniteReadLimit() { cr.remain = maxInt64 }
=== 1700
w.WriteHeader(StatusOK)
}
w.w.Flush()
putBufioWriter(w.w)
w.cw.close()
w.conn.bufw.Flush()
w.conn.r.abortPendingRead()
w.reqBody.registerOnHitEOF(nil) // prevent new background read from starting
if w.canWriteContinue.Load() {
w.disableWriteContinue(true)
}Bash grep -n "handleReadErrorLocked" -A 20 server.go | head -30
grep -n "handleReadErrorLocked" -A 20 server.go | head -30
Read handleReadErrorLocked
777: cr.handleReadErrorLocked(err)
778- }
779- cr.aborted = false
780- cr.inRead = false
781- cr.unlock()
782- cr.cond.Broadcast()
783-}
784-
785-func (cr *connReader) abortPendingRead() {
786- cr.lock()
787- defer cr.unlock()
788- if !cr.inRead {
789- return
790- }
791- cr.aborted = true
792- cr.rwc.SetReadDeadline(aLongTimeAgo)
793- for cr.inRead {
794- cr.cond.Wait()
795- }
796- cr.rwc.SetReadDeadline(time.Time{})
797-}
--
803:// handleReadErrorLocked is called whenever a Read from the client returns a
804-// non-nil error.
805-//
806-// The provided non-nil err is almost always io.EOF or a "use of
807-// closed network connection". In any case, the error is not
808-// particularly interesting, except perhaps for debugging during
809-// development. Any error means the connection is dead and we should
810-// down its context.Bash sed -n '800,820p' server.go
sed -n '800,820p' server.go
Full handleReadErrorLocked
func (cr *connReader) setInfiniteReadLimit() { cr.remain = maxInt64 }
func (cr *connReader) hitReadLimit() bool { return cr.remain <= 0 }
// handleReadErrorLocked is called whenever a Read from the client returns a
// non-nil error.
//
// The provided non-nil err is almost always io.EOF or a "use of
// closed network connection". In any case, the error is not
// particularly interesting, except perhaps for debugging during
// development. Any error means the connection is dead and we should
// down its context.
//
// The caller must hold connReader.mu.
func (cr *connReader) handleReadErrorLocked(_ error) {
if cr.conn == nil {
return
}
cr.conn.cancelCtx()
if res := cr.conn.curReq.Load(); res != nil {
res.closeNotify()
}Bash cd /tmp; for f in nethttp.txt s_*.txt; do echo "== $f"; python3 - "$f" <<'EOF' ...
cd /tmp; for f in nethttp.txt s_*.txt; do echo "== $f"; python3 - "$f" <<'EOF'
import sys
s=open(sys.argv[1]).read(); i=s.index('panic: test timed out'); s=s[i:]
for pat in ['backgroundRead','persistConn).readLoop','persistConn).writeLoop','FD).Read','conn).serve']:
print(' %-24s %d' % (pat, s.count(pat)))
EOF
doneCheck for background read and client goroutines
== nethttp.txt backgroundRead 0 persistConn).readLoop 0 persistConn).writeLoop 0 FD).Read 0 conn).serve 1 == s_8675856801014264801.txt backgroundRead 0 persistConn).readLoop 0 persistConn).writeLoop 0 FD).Read 0 conn).serve 1 == s_8676093236579846177.txt backgroundRead 0 persistConn).readLoop 0 persistConn).writeLoop 0 FD).Read 0 conn).serve 1 == s_8681022390813019185.txt backgroundRead 0 persistConn).readLoop 0 persistConn).writeLoop 0 FD).Read 0 conn).serve 1 == s_8686112892523717025.txt backgroundRead 0 persistConn).readLoop 0 persistConn).writeLoop 0 FD).Read 0 conn).serve 1 Shell cwd was reset to /home/claude/golang_go_78438
Bash sed -n '2112,2126p' /home/claude/golang_go_78438/go/src/net/http/server.go; ls /home/claude/golang_go_78438/go/bin 2>/de...
sed -n '2112,2126p' /home/claude/golang_go_78438/go/src/net/http/server.go; ls /home/claude/golang_go_78438/go/bin 2>/dev/null; /home/claude/golang_go_78438/go/bin/go version 2>&1 | head -2
Confirm startBackgroundRead call and check toolchain
}
c.curReq.Store(w)
// Start background read, which detects when a client has closed its connection
// while a request handler is still running. When the request has a body, we
// start the background read only after the entire body has been consumed.
if w.reqBody.bodyRemains() {
w.reqBody.registerOnHitEOF(w.conn.r.startBackgroundRead)
} else {
w.conn.r.startBackgroundRead()
}
// HTTP cannot have multiple simultaneous active requests.[*]
// Until the server replies to this request, it can't read another,
go
gofmt
go version go1.28-devel_13b477edc7 Thu Aug 20 11:16:25 2026 -0400 linux/amd64Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; mkdir -p $SP; cat > $S...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; mkdir -p $SP; cat > $SP/repro.go <<'EOF'
// Minimal reproducer for the hang seen in net/http.TestTransportGzip/h1 on
// solaris-amd64 (go.dev/issue/78438).
//
// It mimics what net/http does on the server side of a loopback connection:
// a background goroutine blocked in Read on the connection (the server's
// "background read", which detects a client that goes away), while the
// handler goroutine is blocked in a large Write.
//
// The client reads a little of the response and then closes the connection
// with data still unread, which should reset the connection and make the
// server's blocked Write fail.
//
// On Linux the Write fails immediately. On solaris-amd64 the failing builds
// show the Read returning an error while the Write stays blocked forever.
package main
import (
"fmt"
"io"
"net"
"os"
"time"
)
const writeSize = 1 << 20
func main() {
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
panic(err)
}
defer ln.Close()
readDone := make(chan string, 1)
writeDone := make(chan string, 1)
go func() {
c, err := ln.Accept()
if err != nil {
panic(err)
}
// Mimic net/http's connReader.backgroundRead.
go func() {
var b [1]byte
n, err := c.Read(b[:])
readDone <- fmt.Sprintf("Read returned n=%d err=%v", n, err)
}()
// Mimic a handler writing a large response body.
start := time.Now()
n, err := c.Write(make([]byte, writeSize))
writeDone <- fmt.Sprintf("Write returned n=%d err=%v after %v", n, err, time.Since(start).Round(time.Millisecond))
}()
c, err := net.Dial("tcp", ln.Addr().String())
if err != nil {
panic(err)
}
// Read a little, then close with the rest unread.
if _, err := io.ReadFull(c, make([]byte, 4096)); err != nil {
panic(err)
}
time.Sleep(100 * time.Millisecond) // let the server block in Write
if err := c.Close(); err != nil {
panic(err)
}
fmt.Println("client closed the connection")
timeout := time.After(30 * time.Second)
for i := 0; i < 2; i++ {
select {
case s := <-readDone:
fmt.Println("server:", s)
case s := <-writeDone:
fmt.Println("server:", s)
case <-timeout:
fmt.Println("FAIL: server side still blocked 30s after the client closed")
os.Exit(1)
}
}
fmt.Println("PASS")
}
EOF
cd $SP && /home/claude/golang_go_78438/go/bin/go run repro.goWrite and run reproducer on Linux
client closed the connection server: Write returned n=1048576 err=<nil> after 0s server: Read returned n=0 err=read tcp 127.0.0.1:41037->127.0.0.1:43202: read: connection reset by peer PASS Shell cwd was reset to /home/claude/golang_go_78438
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad ...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad
python3 - "$SP/repro.go" <<'EOF'
import sys
p=sys.argv[1]; s=open(p).read()
s=s.replace(''' c, err := ln.Accept()
if err != nil {
panic(err)
}''',''' c, err := ln.Accept()
if err != nil {
panic(err)
}
// Keep the socket buffers small so that the write below actually
// blocks; on the builder it blocks after ~4kB.
c.(*net.TCPConn).SetWriteBuffer(4096)''')
s=s.replace(''' c, err := net.Dial("tcp", ln.Addr().String())
if err != nil {
panic(err)
}''',''' c, err := net.Dial("tcp", ln.Addr().String())
if err != nil {
panic(err)
}
c.(*net.TCPConn).SetReadBuffer(4096)''')
open(p,'w').write(s)
EOF
cd $SP && /home/claude/golang_go_78438/go/bin/go run repro.goMake the write actually block and rerun
client closed the connection server: Write returned n=65536 err=write tcp 127.0.0.1:43727->127.0.0.1:53658: write: connection reset by peer after 101ms server: Read returned n=0 err=EOF PASS Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /home/claude/golang_go_78438/go && git log --oneline --since=2025-11-01 --until=2026-04-15 -- src/net/http/transport_...
cd /home/claude/golang_go_78438/go && git log --oneline --since=2025-11-01 --until=2026-04-15 -- src/net/http/transport_test.go | head -20; echo "=== blame region"; git log -1 --format='%h %ad %s' --date=short $(git log -1 --format=%H -- src/net/http/transport_test.go) >/dev/null; git blame -L 1219,1235 --date=short -s src/net/http/transport_test.go 2>/dev/null | head -20
Check recent history of the test
ccf4ba99a9 net/http: run tests for HTTP/3 where it can already pass
080aa8e964 net/http: use net/http/internal/http2 rather than h2_bundle.go
cbb37e310e net/http: fix WaitGroup race in TestTransportNoReuseAfterEarlyResponse
11d5284363 net/http: try to drain response body upon closing for better connection re-use
=== blame region
6e71e1ca766 src/pkg/http/transport_test.go 1219) var w io.Writer = rw
6e71e1ca766 src/pkg/http/transport_test.go 1220) var buf bytes.Buffer
6e71e1ca766 src/pkg/http/transport_test.go 1221) if req.FormValue("chunked") == "0" {
6e71e1ca766 src/pkg/http/transport_test.go 1222) w = &buf
6e71e1ca766 src/pkg/http/transport_test.go 1223) defer io.Copy(rw, &buf)
6e71e1ca766 src/pkg/http/transport_test.go 1224) defer func() {
6e71e1ca766 src/pkg/http/transport_test.go 1225) rw.Header().Set("Content-Length", strconv.Itoa(buf.Len()))
6e71e1ca766 src/pkg/http/transport_test.go 1226) }()
6e71e1ca766 src/pkg/http/transport_test.go 1227) }
cc9ed447d0a src/pkg/net/http/transport_test.go 1228) gz := gzip.NewWriter(w)
92210eefb2f src/pkg/http/transport_test.go 1229) gz.Write([]byte(testString))
6e71e1ca766 src/pkg/http/transport_test.go 1230) if req.FormValue("body") == "large" {
39493be2948 src/pkg/http/transport_test.go 1231) io.CopyN(gz, rand.Reader, nRandBytes)
6e71e1ca766 src/pkg/http/transport_test.go 1232) }
6e71e1ca766 src/pkg/http/transport_test.go 1233) gz.Close()
747e1961e95 src/net/http/transport_test.go 1234) })).ts
3b988eb643e src/net/http/transport_test.go 1235) c := ts.Client()Bash git log -1 --format='%H%n%ad%n%s%n%n%b' --date=short 11d5284363 | head -50
git log -1 --format='%H%n%ad%n%s%n%n%b' --date=short 11d5284363 | head -50
Inspect drain-on-close commit
11d5284363ed88b8cc8ec6e68db80a16b2f9e708 2026-01-21 net/http: try to drain response body upon closing for better connection re-use Currently, we have a rather inconsistent behavior in terms of whether a connection can be re-used or not when an HTTP body is not read to completion: - In HTTP/2, not reading bodies to completion is not an issue, since a new HTTP/2 stream can be created on the same TCP connection. - In HTTP/1 server, we discard up to 256 KiB of unconsumed request body, to potentially allow re-use. - In HTTP/1 client, we do not do anything, and fail to re-use a TCP connection if there are any unconsumed response body at all. This has led to some confusion. For example, some users have mistakenly discarded response body for HTTP/2 when doing so is not needed. Manually discarding response body can also be disadvantageous if the body is excessively large or is a never-ending stream. To solve this issue, this CL makes it so that closing a response body will cause any remaining content to be drained, up to a limit of 256 KiB or 50 milliseconds, whichever one is reached first. This allows better connection re-use for HTTP/1, and most users can now avoid having to manually drain their response body. For #77370 Change-Id: I71e1227fc9cf5f901362c8e234320817f6b0be24 Reviewed-on: https://go-review.googlesource.com/c/go/+/737720 Reviewed-by: Nicholas Husin <husin@google.com> LUCI-TryBot-Result: Go LUCI <golang-scoped@luci-project-accounts.iam.gserviceaccount.com> Reviewed-by: Damien Neil <dneil@google.com>
Bash cd /home/claude/golang_go_78438/go/src/net/http && grep -n "drain\|Drain" transport.go | head -40
cd /home/claude/golang_go_78438/go/src/net/http && grep -n "drain\|Drain" transport.go | head -40
Find drain implementation
2437:// read when draining the response body of any unread bytes after it has been
2442:// willing to spend on draining a response body of any unread bytes after it
2446:func maybeDrainBody(body io.Reader) bool {
2447: drainedCh := make(chan bool, 1)
2450: drainedCh <- true
2452: drainedCh <- false
2456: case drained := <-drainedCh:
2457: return drained
2624: tryDrain := !bodyEOF && resp.ContentLength <= maxPostCloseReadBytes
2625: if tryDrain {
2627: bodyEOF = maybeDrainBody(body.body)
2634: if !tryDrain && bodyEOF {Bash sed -n '2430,2470p' transport.go; echo "=== readLoop region"; sed -n '2600,2680p' transport.go
sed -n '2430,2470p' transport.go; echo "=== readLoop region"; sed -n '2600,2680p' transport.go
Read drain code and readLoop handling
// errCallerOwnsConn is an internal sentinel error used when we hand
// off a writable response.Body to the caller. We use this to prevent
// closing a net.Conn that is now owned by the caller.
var errCallerOwnsConn = errors.New("read loop ending; caller owns writable underlying conn")
// maxPostCloseReadBytes is the max number of bytes that a client is willing to
// read when draining the response body of any unread bytes after it has been
// closed. This number is chosen for consistency with maxPostHandlerReadBytes.
const maxPostCloseReadBytes = 256 << 10
// maxPostCloseReadTime defines the maximum amount of time that a client is
// willing to spend on draining a response body of any unread bytes after it
// has been closed.
const maxPostCloseReadTime = 50 * time.Millisecond
func maybeDrainBody(body io.Reader) bool {
drainedCh := make(chan bool, 1)
go func() {
if _, err := io.CopyN(io.Discard, body, maxPostCloseReadBytes+1); err == io.EOF {
drainedCh <- true
} else {
drainedCh <- false
}
}()
select {
case drained := <-drainedCh:
return drained
case <-time.After(maxPostCloseReadTime):
return false
}
}
func (pc *persistConn) readLoop() {
closeErr := errReadLoopExiting // default value, if not changed below
defer func() {
pc.close(closeErr)
pc.t.removeIdleConn(pc)
if pc.internalStateHook != nil {
pc.internalStateHook()
}
=== readLoop region
return err
},
}
resp.Body = body
if rc.addedGzip && ascii.EqualFold(resp.Header.Get("Content-Encoding"), "gzip") {
resp.Body = &gzipReader{body: body}
resp.Header.Del("Content-Encoding")
resp.Header.Del("Content-Length")
resp.ContentLength = -1
resp.Uncompressed = true
}
select {
case rc.ch <- responseAndError{res: resp}:
case <-rc.callerGone:
return
}
// Before looping back to the top of this function and peeking on
// the bufio.Reader, wait for the caller goroutine to finish
// reading the response body. (or for cancellation or death)
select {
case bodyEOF := <-waitForBodyRead:
tryDrain := !bodyEOF && resp.ContentLength <= maxPostCloseReadBytes
if tryDrain {
eofc <- struct{}{}
bodyEOF = maybeDrainBody(body.body)
}
alive = alive &&
bodyEOF &&
!pc.sawEOF &&
pc.wroteRequest() &&
tryPutIdleConn(rc.treq)
if !tryDrain && bodyEOF {
eofc <- struct{}{}
}
case <-rc.treq.ctx.Done():
alive = false
pc.cancelRequest(context.Cause(rc.treq.ctx))
case <-pc.closech:
alive = false
}
rc.treq.cancel(errRequestDone)
testHookReadLoopBeforeNextRead()
}
}
func (pc *persistConn) readLoopPeekFailLocked(peekErr error) {
if pc.closed != nil {
return
}
if n := pc.br.Buffered(); n > 0 {
buf, _ := pc.br.Peek(n)
if is408Message(buf) {
pc.closeLocked(errServerClosedIdle)
return
} else {
log.Printf("Unsolicited response received on idle HTTP channel starting with %q; err=%v", buf, peekErr)
}
}
if peekErr == io.EOF {
// common case.
pc.closeLocked(errServerClosedIdle)
} else {
pc.closeLocked(fmt.Errorf("readLoopPeekFailLocked: %w", peekErr))
}
}
// is408Message reports whether buf has the prefix of an
// HTTP 408 Request Timeout response.
// See golang.org/issue/32310.
func is408Message(buf []byte) bool {
if len(buf) < len("HTTP/1.x 408") {
return false
}
if string(buf[:7]) != "HTTP/1." {
return false
}
return string(buf[8:12]) == " 408"I've gone through all six failure logs. Here's the triage.
What the watchflakes snippet hides
The issue body is useless on its own — every entry is [SKIP] with just === RUN/PAUSE/CONT. The real failure in each case is a package-level timeout of net/http, with TestTransportGzip/h1 named as the last test still running:
panic: test timed out after 10m0s
running tests:
TestTransportGzip/h1 (9m58s)
So TestTransportGzip isn't failing an assertion — it's wedged.
The dominant signature: 5 of 6 failures, all solaris-amd64
All five solaris logs (2026-03-28, 05-19, 07-13, 07-15, 08-21) are byte-for-byte the same shape:
httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0x... 127.0.0.1:47089 in state active
and exactly two relevant goroutines left alive:
-
The server handler, parked in netpoll for the full 9–17 minutes:
internal/poll.(*FD).Write({0x…f97, 0xff10d, 0x11f069}) net/http.(*chunkWriter).Write({0x…000, 0x1000a4, 0x120000}) net/http_test.testTransportGzip.func1.deferwrap1() transport_test.go:1223That's the
defer io.Copy(rw, &buf)on thechunked=0path — writing the buffered ~1 MB body. The offsets are identical in every log (0xf97= 3991): the server's 4 KiBbufioflushed 105 bytes of header + 3991 bytes of body fine, and the second, direct 1 MB write is the one that blocks. -
The test cleanup, in
httptest.Server.Close→wg.Wait(). The test body itself completed successfully;Server.Closedeliberately never force-closesStateActiveconns, so it waits forever.
The client side is completely gone from the dump — no persistConn.readLoop/writeLoop. The transport did close its socket (Body.Close blocks on <-eofc until readLoop runs pc.close, and the test moved on). So: the server is blocked writing to a TCP connection whose peer is closed, and the write never returns.
The sharpest clue: net/http starts a connReader.backgroundRead goroutine on the same fd for a body-less request (server.go:2122), and it is absent from all five dumps — it already returned. Its only exit for this connection is a read error, which also fires handleReadErrorLocked → cancelCtx. So on the same fd, at the same moment:
- the read side got the connection error and woke,
- the write side, armed for
POLLOUTviaport_associate, was never woken.
That points at Solaris event ports / the OS not reporting write-side readiness or POLLERR for a reset socket, not at anything in net/http. Note netpoll_solaris.go only wakes a writer on POLLOUT|POLLHUP|POLLERR; a bare POLLIN event re-arms POLLOUT and nothing more ever arrives.
I wrote a standalone reproducer that isolates exactly this (background reader + blocked 1 MB write + peer closes with data unread), no net/http involved:
…/scratchpad/repro.go — on linux/amd64 it passes:
server: Write returned n=65536 err=write tcp …: connection reset by peer after 101ms
server: Read returned n=0 err=EOF
PASS
It's worth running on a solaris gomote; I'd expect it to print FAIL: server side still blocked.
Likely trigger for why this started in March
11d5284363 (2026-01-21, "net/http: try to drain response body upon closing for better connection re-use") changed what res.Body.Close() does. Previously the transport closed the socket immediately with a full receive buffer → guaranteed RST → server write fails. Now maybeDrainBody reads up to 256 KiB / 50 ms first, so the socket is often closed with an empty receive buffer → FIN instead of RST, and the race window differs. (For this test tryDrain is always taken: the gzip path sets resp.ContentLength = -1, and -1 <= maxPostCloseReadBytes.) That's consistent with the first flake appearing after that CL, and with it being rare.
The ppc64le outlier is a different failure
gotip-linux-ppc64le_power8 (2026-05-18) has nothing in common with the solaris five. No blocked handler, no blocked in Close. The test goroutine is stuck 17 minutes in Transport.getConn's select, and its dial goroutine — created by that same getConn 17 minutes earlier — is [runnable], preempted at net.(*Dialer).deadline, i.e. it had barely begun dialing at panic time. That reads like a scheduler wakeup/starvation problem, not a net/http one, and I'd treat it as unrelated noise on this issue rather than evidence about TestTransportGzip.
Suggested disposition
- Not a
net/httpbug and not a duplicate of anet/httpissue. JunyangShao's "just timed out?" is right about the symptom; the cause is the wedged handler write. - It belongs with the Solaris umbrella, #60152 (
all: test timeouts on the oraclerel builder, label OS-Solaris), cc@golang/solaris/@rorth, with the reproducer above. - If you want the test to stop wedging CI regardless, the cheap fix is a
ts.CloseClientConnections()before cleanup intestTransportGzip(it force-closes the server side of the conn, which unblocks the write); the honest fix is at the OS/netpoll layer. - The ppc64le entry should probably be split into its own issue.
Want me to draft the comment for the issue thread, or go deeper on the ppc64le scheduler angle?
You Wed 26 Aug 2026 11:52:16
The ppc64le one is probably #78576. Lets focus on Solaris. You can create a gomote with 'gomote create gotip-solaris-amd64' if you need to try something.
Assistant Wed 26 Aug 2026 11:52:21
Ran tools (6 Bash)
Bash which gomote; gomote version 2>&1 | head -3; gomote group ls 2>&1 | head -5
which gomote; gomote version 2>&1 | head -3; gomote group ls 2>&1 | head -5
Check gomote availability
/home/claude/bin/gomote # Unknown command "version" Usage of gomote: gomote [global-flags] <cmd> [cmd-flags] # Error running group: unknown sub-command "ls"
Bash gomote list 2>&1 | head; echo "---help---"; gomote 2>&1 | head -40
gomote list 2>&1 | head; echo "---help---"; gomote 2>&1 | head -40
List gomotes and show help
rsc-gotip-windows-386-0 gotip-windows-386 gotip-windows-386 expires in 29m59.011521886s
rsc-gotip-windows-386-1 gotip-windows-386 gotip-windows-386 expires in 29m59.011493833s
rsc-gotip-windows-386-2 gotip-windows-386 gotip-windows-386 expires in 29m57.011491488s
---help---
Usage of gomote: gomote [global-flags] <cmd> [cmd-flags]
Global flags:
-group string
name of the gomote group to apply commands to (default is $GOMOTE_GROUP)
-server string
Address for GRPC server (default "gomote.golang.org:443")
Commands:
create create a buildlet; with no args, list types of buildlets
destroy destroy a buildlet
gettar extract a tar.gz from a buildlet
group manage groups of instances
list list active buildlets
login authenticate with the gomote service
ls list the contents of a directory on a buildlet
ping test whether a buildlet is alive and reachable
push sync your GOROOT directory to the buildlet
put put files on a buildlet
puttar extract a tar.gz to a buildlet
rdp Unimplimented: RDP (Remote Desktop Protocol) to a Windows buildlet
repro reproduce a build environment in a new buildlet
rm delete files or directories
run run a command on a buildlet
ssh ssh to a buildletBash gomote create gotip-solaris-amd64 2>&1 | tail -5
gomote create gotip-solaris-amd64 2>&1 | tail -5
Create solaris gomote
600000
# still creating gotip-solaris-amd64 (1) after 41s; 0 requests ahead of you # still creating gotip-solaris-amd64 (1) after 46s; 0 requests ahead of you # still creating gotip-solaris-amd64 (1) after 51s; 0 requests ahead of you # still creating gotip-solaris-amd64 (1) after 56s; 0 requests ahead of you rsc-gotip-solaris-amd64-0
Bash G=rsc-gotip-solaris-amd64-0; gomote ls $G 2>&1 | head -20; echo "---"; gomote run $G /bin/sh -c 'uname -a; echo GOROOT=$...
G=rsc-gotip-solaris-amd64-0; gomote ls $G 2>&1 | head -20; echo "---"; gomote run $G /bin/sh -c 'uname -a; echo GOROOT=$GOROOT; ls /' 2>&1 | head -20
Inspect gomote environment
dgrwxr-xr-x .cipd/ -rwxr--r-- buildlet 23394567 2026-08-26T15:53:30Z dgrwxr-xr-x cache/ dgrwxr-xr-x gocache/ dgrwxr-xr-x goplscache/ dgrwxr-xr-x tmp/ dgrwxr-xr-x tools/ --- # Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote1892719975/rsc-gotip-solaris-amd64-0.stdout"... SunOS s11-i386.foss 5.11 11.4.95.226.0 i86pc i386 i86pc kernel-zone GOROOT= bin boot dev devices etc export home kernel lib media mnt opt platform proc root rpool sbin
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/solrepro.go ...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/solrepro.go <<'EOF'
// Reproducer for go.dev/issue/78438: on solaris-amd64, a server goroutine
// blocked writing to a TCP connection is apparently never woken when the peer
// closes the connection, even though a concurrent Read on the same fd does see
// the error.
//
// This mimics what net/http does: connReader.backgroundRead sits in Read on the
// connection while the handler writes the response body.
package main
import (
"flag"
"fmt"
"io"
"net"
"os"
"time"
)
var (
writeSize = flag.Int("writesize", 1<<20, "bytes the server writes")
sndbuf = flag.Int("sndbuf", 0, "SO_SNDBUF on the server conn (0: default)")
rcvbuf = flag.Int("rcvbuf", 0, "SO_RCVBUF on the client conn (0: default)")
wait = flag.Duration("wait", 20*time.Second, "how long to wait for the server side")
runs = flag.Int("runs", 5, "iterations per mode")
)
type result struct {
who string
text string
}
// one runs a single iteration. drain reports whether the client drains the way
// net/http's Transport does since CL 737720 (up to 256KiB / 50ms) before
// closing; otherwise it closes immediately with the data unread.
func one(drain bool) bool {
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
panic(err)
}
defer ln.Close()
res := make(chan result, 2)
go func() {
c, err := ln.Accept()
if err != nil {
panic(err)
}
if *sndbuf > 0 {
c.(*net.TCPConn).SetWriteBuffer(*sndbuf)
}
// net/http's connReader.backgroundRead.
go func() {
var b [1]byte
n, err := c.Read(b[:])
res <- result{"Read", fmt.Sprintf("n=%d err=%v", n, err)}
}()
start := time.Now()
n, err := c.Write(make([]byte, *writeSize))
res <- result{"Write", fmt.Sprintf("n=%d err=%v after %v", n, err, time.Since(start).Round(time.Millisecond))}
}()
c, err := net.Dial("tcp", ln.Addr().String())
if err != nil {
panic(err)
}
if *rcvbuf > 0 {
c.(*net.TCPConn).SetReadBuffer(*rcvbuf)
}
// Read a little of the response, like the test does.
if _, err := io.ReadFull(c, make([]byte, 4096)); err != nil {
panic(err)
}
time.Sleep(100 * time.Millisecond) // let the server block in Write
if drain {
done := make(chan bool, 1)
go func() {
_, err := io.CopyN(io.Discard, c, 256<<10+1)
done <- err == io.EOF
}()
select {
case <-done:
case <-time.After(50 * time.Millisecond):
}
}
c.Close()
ok := true
deadline := time.After(*wait)
got := map[string]bool{}
for i := 0; i < 2; i++ {
select {
case r := <-res:
got[r.who] = true
fmt.Printf(" %-5s %s\n", r.who, r.text)
case <-deadline:
for _, who := range []string{"Read", "Write"} {
if !got[who] {
fmt.Printf(" %-5s STUCK: still blocked %v after the client closed\n", who, *wait)
}
}
ok = false
i = 2
}
}
return ok
}
func main() {
flag.Parse()
fmt.Printf("%s/%s writesize=%d sndbuf=%d rcvbuf=%d\n", os.Getenv("GOOS"), os.Getenv("GOARCH"), *writeSize, *sndbuf, *rcvbuf)
bad := 0
for _, drain := range []bool{false, true} {
name := "close immediately (pre-CL 737720)"
if drain {
name = "drain 256KiB/50ms, then close (post-CL 737720)"
}
fmt.Printf("== %s\n", name)
for i := 0; i < *runs; i++ {
fmt.Printf(" run %d:\n", i)
if !one(drain) {
bad++
}
}
}
if bad > 0 {
fmt.Printf("FAIL: %d stuck\n", bad)
os.Exit(1)
}
fmt.Println("PASS")
}
EOF
cd $SP && GOOS=solaris GOARCH=amd64 /home/claude/golang_go_78438/go/bin/go build -o solrepro.solaris solrepro.go && ls -l solrepro.solaris && /home/claude/golang_go_78438/go/bin/go run solrepro.go -runs=2 -wait=5s 2>&1 | tail -20Build solaris repro and sanity-check on Linux
-rwxrwxr-x 1 claude claude 3437998 Aug 26 11:54 solrepro.solaris
/ writesize=1048576 sndbuf=0 rcvbuf=0
== close immediately (pre-CL 737720)
run 0:
Write n=1048576 err=<nil> after 0s
Read n=0 err=read tcp 127.0.0.1:33313->127.0.0.1:59340: read: connection reset by peer
run 1:
Write n=1048576 err=<nil> after 0s
Read n=0 err=read tcp 127.0.0.1:36893->127.0.0.1:50062: read: connection reset by peer
== drain 256KiB/50ms, then close (post-CL 737720)
run 0:
Write n=1048576 err=<nil> after 1ms
Read n=0 err=read tcp 127.0.0.1:41403->127.0.0.1:34116: read: connection reset by peer
run 1:
Write n=1048576 err=<nil> after 1ms
Read n=0 err=read tcp 127.0.0.1:43237->127.0.0.1:46654: read: connection reset by peer
PASS
Shell cwd was reset to /home/claude/golang_go_78438Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/solrepro.solaris solrepro 2>&1 | tail -3; gomote run $G /bin/sh -c 'chmod +x ./solrepro && ./solrepro -runs=3 -wait=15s' 2>&1 | tail -40
Run reproducer on solaris gomote
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote927603088/rsc-gotip-solaris-amd64-0.stdout"...
solaris/amd64 writesize=1048576 sndbuf=0 rcvbuf=0
== close immediately (pre-CL 737720)
run 0:
Read n=0 err=EOF
Write n=294912 err=write tcp 127.0.0.1:60024->127.0.0.1:48934: write: broken pipe after 101ms
run 1:
Read n=0 err=EOF
Write n=294912 err=write tcp 127.0.0.1:45546->127.0.0.1:37628: write: broken pipe after 101ms
run 2:
Read n=0 err=EOF
Write n=294912 err=write tcp 127.0.0.1:36328->127.0.0.1:55648: write: broken pipe after 100ms
== drain 256KiB/50ms, then close (post-CL 737720)
run 0:
Read n=0 err=EOF
Write n=294912 err=write tcp 127.0.0.1:61518->127.0.0.1:46457: write: broken pipe after 103ms
run 1:
Read n=0 err=EOF
Write n=270336 err=write tcp 127.0.0.1:34941->127.0.0.1:60290: write: broken pipe after 102ms
run 2:
Read n=0 err=read tcp 127.0.0.1:58348->127.0.0.1:48244: read: connection reset by peer
Write n=294912 err=write tcp 127.0.0.1:58348->127.0.0.1:48244: write: broken pipe after 102ms
PASS
# Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote927603088/rsc-gotip-solaris-amd64-0.stdout".The simple repro doesn't hang on Solaris — writes get EPIPE promptly. So my OS-level hypothesis is wrong as stated. Let me run the real test there.
Ran tools (4 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go/src && GOOS=solaris GOARCH=amd64 /home/claude/golang_go_78438/go/bin/go test -c -o $SP/nethttp.test net/http 2>&1 | tail -5; ls -l $SP/nethttp.test
Cross-compile net/http test binary for solaris
300000
-rwxrwxr-x 1 claude claude 17900891 Aug 26 11:56 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/nethttp.test
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/nethttp.test nethttp.test 2>&1 | tail -2; gomote run $G /bin/sh -c 'chmod +x nethttp.test && ./nethttp.test -test.run "^TestTransportGzip$" -test.count=200 -test.timeout=4m 2>&1 | tail -20' 2>&1 | tail -25
Stress TestTransportGzip on solaris
400000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote1619022165/rsc-gotip-solaris-amd64-0.stdout"... PASS # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote1619022165/rsc-gotip-solaris-amd64-0.stdout".
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'psrinfo | wc -l; echo ---; for i in 1 2 3 4 5 6; do ./nethttp.tes...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'psrinfo | wc -l; echo ---; for i in 1 2 3 4 5 6; do ./nethttp.test -test.run "^TestTransportGzip$" -test.count=400 -test.timeout=6m > out.$i 2>&1 & done; wait; grep -l -v "^PASS" out.* 2>/dev/null; for i in 1 2 3 4 5 6; do echo "run $i: $(tail -1 out.$i)"; done' 2>&1 | tail -20
Run 6 concurrent stress instances
600000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote235507741/rsc-gotip-solaris-amd64-0.stdout"... 21 ---
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'ls -l out.*; for i in 1 2 3 4 5 6; do echo "== run $i: $(tail -2 ...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'ls -l out.*; for i in 1 2 3 4 5 6; do echo "== run $i: $(tail -2 out.$i | tr "\n" " ")"; done' 2>&1 | tail -20
Check stress results
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote158807008/rsc-gotip-solaris-amd64-0.stdout"... -rw-r--r-- 1 swarming golang 5 Aug 26 17:57 out.1 -rw-r--r-- 1 swarming golang 5 Aug 26 17:57 out.2 -rw-r--r-- 1 swarming golang 6003 Aug 26 18:02 out.3 -rw-r--r-- 1 swarming golang 5 Aug 26 17:57 out.4 -rw-r--r-- 1 swarming golang 5994 Aug 26 18:02 out.5 -rw-r--r-- 1 swarming golang 6003 Aug 26 18:02 out.6 == run 1: PASS == run 2: PASS == run 3: created by testing.(*T).Run in goroutine 4876 /home/claude/golang_go_78438/go/src/testing/testing.go:2265 +0x4d4 == run 4: PASS == run 5: created by net/http.(*Server).Serve in goroutine 434 /home/claude/golang_go_78438/go/src/net/http/server.go:3593 +0x4cc == run 6: created by net/http.(*Server).Serve in goroutine 21233 /home/claude/golang_go_78438/go/src/net/http/server.go:3593 +0x4cc # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote158807008/rsc-gotip-solaris-amd64-0.stdout".
Reproduced — 3 of 6 concurrent stress instances hung.
Ran tools (1 Bash)
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'cat out.3' 2>&1 | head -80
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'cat out.3' 2>&1 | head -80
Fetch hung run output
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2187383444/rsc-gotip-solaris-amd64-0.stdout"...
2026/08/26 17:57:09 httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0x35aeae7080a8 127.0.0.1:48106 in state active
panic: test timed out after 6m0s
running tests:
TestTransportGzip (5m55s)
TestTransportGzip/h1 (5m55s)
goroutine 4907 [running]:
testing.(*M).startAlarm.func1()
/home/claude/golang_go_78438/go/src/testing/testing.go:2966 +0x34a
created by time.goFunc
/home/claude/golang_go_78438/go/src/time/sleep.go:187 +0x2d
goroutine 1 [chan receive, 5 minutes]:
testing.(*T).Run(0x35aeae740488, {0x9a1c8c?, 0x35aeae7d1a70?}, 0xf57880)
/home/claude/golang_go_78438/go/src/testing/testing.go:2273 +0x4f2
testing.runTests.func1(0x35aeae740488)
/home/claude/golang_go_78438/go/src/testing/testing.go:2749 +0x37
testing.tRunner(0x35aeae740488, 0x35aeae7d1b98)
/home/claude/golang_go_78438/go/src/testing/testing.go:2200 +0xea
testing.runTests({0x0, 0x0}, {0x99a03a, 0x8}, 0x35aeae7221c8, {0xfd1550, 0x26b, 0x26b}, {0xc29be16ce359429c, 0x53d1d64a38, ...})
/home/claude/golang_go_78438/go/src/testing/testing.go:2747 +0x4d0
testing.(*M).Run(0x35aeae73a3c0)
/home/claude/golang_go_78438/go/src/testing/testing.go:2607 +0x68f
net/http_test.TestMain(0xff88e0?)
/home/claude/golang_go_78438/go/src/net/http/main_test.go:24 +0x27
main.main()
_testmain.go:1354 +0xa5
goroutine 4916 [IO wait, 5 minutes]:
internal/poll.runtime_pollWait(0x7fffbf538c00, 0x77)
/home/claude/golang_go_78438/go/src/runtime/netpoll.go:351 +0x8b
internal/poll.(*pollDesc).wait(0x35aeae736480?, 0x35aeaf046f97?, 0x0)
/home/claude/golang_go_78438/go/src/internal/poll/fd_poll_runtime.go:84 +0x27
internal/poll.(*pollDesc).waitWrite(...)
/home/claude/golang_go_78438/go/src/internal/poll/fd_poll_runtime.go:93
internal/poll.(*FD).Write(0x35aeae736480, {0x35aeaefb0f97, 0xff10d, 0x11f069})
/home/claude/golang_go_78438/go/src/internal/poll/fd_unix.go:395 +0x30a
net.(*netFD).Write(0x35aeae736480, {0x35aeaefb0f97?, 0x0?, 0x0?})
/home/claude/golang_go_78438/go/src/net/fd_posix.go:109 +0x25
net.(*conn).Write(0x35aeae7080a8, {0x35aeaefb0f97?, 0x35aeae837000?, 0x2?})
/home/claude/golang_go_78438/go/src/net/net.go:208 +0x45
net/http.checkConnErrorWriter.Write({0x35aeaea760c0?}, {0x35aeaefb0f97?, 0xbb15352cbaac2340?, 0xc61bcfb6faa416ae?})
/home/claude/golang_go_78438/go/src/net/http/server.go:4261 +0x26
bufio.(*Writer).Write(0x35aeaea760c0, {0x35aeaefb0000?, 0x1000a4?, 0x120000?})
/home/claude/golang_go_78438/go/src/bufio/bufio.go:682 +0xec
net/http.(*chunkWriter).Write(0x35aeae9ec048, {0x35aeaefb0000, 0x1000a4, 0x120000})
/home/claude/golang_go_78438/go/src/net/http/server.go:392 +0xff
bufio.(*Writer).Write(0x35aeae9ba040, {0x35aeaefb0000?, 0x28?, 0x0?})
/home/claude/golang_go_78438/go/src/bufio/bufio.go:682 +0xec
net/http.(*response).write(0x35aeae9ec000, 0x1000a4, {0x35aeaefb0000, 0x1000a4, 0x120000}, {0x0, 0x0})
/home/claude/golang_go_78438/go/src/net/http/server.go:1688 +0x1be
net/http.(*response).Write(0xe560f8?, {0x35aeaefb0000?, 0xff9e40?, 0x35aeae7cd920?})
/home/claude/golang_go_78438/go/src/net/http/server.go:1651 +0x2a
bytes.(*Buffer).WriteTo(0x35aeaea0c6c0, {0xf4fda8?, 0x35aeae9ec000?})
/home/claude/golang_go_78438/go/src/bytes/buffer.go:279 +0x7b
io.copyBuffer({0xf4fda8, 0x35aeae9ec000}, {0xf4f6c8, 0x35aeaea0c6c0}, {0x0, 0x0, 0x0})
/home/claude/golang_go_78438/go/src/io/io.go:411 +0x9d
io.Copy(...)
/home/claude/golang_go_78438/go/src/io/io.go:388
net/http_test.testTransportGzip.func1({0xf53248, 0x35aeae9ec000}, 0x35aeae9983c0)
/home/claude/golang_go_78438/go/src/net/http/transport_test.go:1234 +0x407
net/http.HandlerFunc.ServeHTTP(0x10?, {0xf53248?, 0x35aeae9ec000?}, 0x0?)
/home/claude/golang_go_78438/go/src/net/http/server.go:2334 +0x29
net/http.serverHandler.ServeHTTP({0x35aeaea76080?}, {0xf53248?, 0x35aeae9ec000?}, 0x1?)
/home/claude/golang_go_78438/go/src/net/http/server.go:3419 +0xbc
net/http.(*conn).serve(0x35aeae9881b0, {0xf541c0, 0x35aeae7c5020})
/home/claude/golang_go_78438/go/src/net/http/server.go:2133 +0x6bc
created by net/http.(*Server).Serve in goroutine 4878
/home/claude/golang_go_78438/go/src/net/http/server.go:3593 +0x4cc
goroutine 4876 [chan receive, 5 minutes]:
testing.(*T).Run(0x35aeae7406c8, {0x99809b?, 0xf5b600?}, 0x35aeae7c4ab0)
/home/claude/golang_go_78438/go/src/testing/testing.go:2273 +0x4f2
net/http_test.run[...](0x35aeae7406c8, 0xf58f48, {0x35aeae6cff50, 0x1, 0x35aeae8cff60})
/home/claude/golang_go_78438/go/src/net/http/clientserver_test.go:141 +0x4b7
net/http_test.TestTransportGzip(0x35aeae7406c8)
/home/claude/golang_go_78438/go/src/net/http/transport_test.go:1203 +0x69
testing.tRunner(0x35aeae7406c8, 0xf57880)Now let me capture OS-level state at the moment of the hang.
Ran tools (2 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/diag.sh <<'E...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/diag.sh <<'EOF'
#!/bin/sh
# Start a batch of stressers, wait for one to wedge, then dump OS state.
rm -f s.*.out
i=0
while [ $i -lt 8 ]; do
./nethttp.test -test.run '^TestTransportGzip$' -test.count=400 -test.timeout=20m > s.$i.out 2>&1 &
echo "$i $!" >> pids.txt
i=`expr $i + 1`
done
n=0
hit=""
while [ $n -lt 100 ]; do
hit=`grep -l "blocked in Close" s.*.out 2>/dev/null | head -1`
if [ -n "$hit" ]; then break; fi
sleep 3
n=`expr $n + 1`
done
if [ -z "$hit" ]; then
echo "NO HANG within timeout"
pkill -f nethttp.test
exit 1
fi
idx=`echo $hit | sed 's/^s\.//; s/\.out$//'`
pid=`grep "^$idx " pids.txt | tail -1 | awk '{print $2}'`
echo "=== wedged: $hit pid=$pid"
grep "in state active" $hit
port=`grep "in state active" $hit | sed 's/.*127\.0\.0\.1://; s/ .*//' | head -1`
echo "=== client port: $port"
echo "=== netstat (all sockets involving that port):"
netstat -an -P tcp -f inet | grep -w "$port"
echo "=== netstat summary of 127.0.0.1 conns for this pid:"
pfiles $pid 2>/dev/null | grep -A3 "SOCK_STREAM" | grep -B1 -A2 "$port"
echo "=== full pfiles entry:"
pfiles $pid 2>/dev/null | sed -n "/$port/,+3p"
echo "=== prstat/truss 3s:"
truss -c -p $pid 2>&1 &
tpid=$!
sleep 3
kill -INT $tpid 2>/dev/null
sleep 1
echo "=== done"
pkill -f nethttp.test
EOF
G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/diag.sh diag.sh 2>&1 | tail -1; gomote run $G /bin/sh -c 'rm -f pids.txt; chmod +x diag.sh; ./diag.sh' 2>&1 | tail -60Capture OS state at hang
600000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote896760144/rsc-gotip-solaris-amd64-0.stdout"...
=== wedged: s.1.out pid=25812
*net.TCPConn 0x6aa5120e000 127.0.0.1:57071 in state active
=== client port: 57071
=== netstat (all sockets involving that port):
=== netstat summary of 127.0.0.1 conns for this pid:
grep: illegal option -- B
grep: illegal option -- 1
grep: illegal option -- A
grep: illegal option -- 2
Usage: grep [-c|-l|-q] -bhinsvw pattern file . . .
=== full pfiles entry:
grep: illegal option -- A
grep: illegal option -- 3
Usage: grep [-c|-l|-q] -bhinsvw pattern file . . .
sed: command garbled: /57071/,+3p
=== prstat/truss 3s:
=== done
signals ------------
SIGTERM 3
total: 3
syscall seconds calls errors
lwp_park .000 1 1
sigaction .000 1
lwp_kill .000 2
lwp_sigmask .000 4
-------- ------ ----
sys totals: .000 8 1
usr time: .000
elapsed: 3.990
# Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote896760144/rsc-gotip-solaris-amd64-0.stdout".Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/diag.sh <<'E...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/diag.sh <<'EOF'
#!/bin/sh
rm -f s.*.out pids.txt
i=0
while [ $i -lt 8 ]; do
./nethttp.test -test.run '^TestTransportGzip$' -test.count=400 -test.timeout=20m > s.$i.out 2>&1 &
echo "$i $!" >> pids.txt
i=`expr $i + 1`
done
n=0
hit=""
while [ $n -lt 100 ]; do
hit=`grep -l "blocked in Close" s.*.out 2>/dev/null | head -1`
if [ -n "$hit" ]; then break; fi
sleep 3
n=`expr $n + 1`
done
if [ -z "$hit" ]; then echo "NO HANG"; pkill -f nethttp.test; exit 1; fi
sleep 5
idx=`echo $hit | sed 's/^s\.//; s/\.out$//'`
pid=`grep "^$idx " pids.txt | tail -1 | awk '{print $2}'`
port=`grep "in state active" $hit | sed 's/.*127\.0\.0\.1://; s/ .*//' | head -1`
echo "=== wedged file=$hit pid=$pid clientport=$port"
echo "=== netstat lines mentioning $port:"
netstat -an -P tcp -f inet > ns.out 2>&1
grep "\.$port " ns.out
grep "\.$port$" ns.out
echo "--- (count of all 127.0.0.1 tcp entries: `grep -c 127.0.0.1 ns.out`)"
echo "=== netstat header + a few loopback lines for format reference:"
head -8 ns.out
grep 127.0.0.1 ns.out | head -5
echo "=== pfiles sockets of the wedged process:"
pfiles $pid > pf.out 2>&1
grep -n "port\|SOCK_STREAM\|sockname\|peername" pf.out | head -60
echo "=== pfiles lines mentioning $port:"
grep -n "$port" pf.out
EOF
G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/diag.sh diag.sh 2>&1 | tail -1; gomote run $G /bin/sh -c 'chmod +x diag.sh; ./diag.sh' 2>&1 | tail -80Better OS-level diagnostics at hang
900000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3657439566/rsc-gotip-solaris-amd64-0.stdout"...
=== wedged file=s.6.out pid=28193 clientport=65381
=== netstat lines mentioning 65381:
--- (count of all 127.0.0.1 tcp entries: 19)
=== netstat header + a few loopback lines for format reference:
TCP: IPv4
Local Address Remote Address Swind Send-Q Rwind Recv-Q State
-------------------- -------------------- ------- ------ ------- ------ -----------
127.0.0.1.5999 *.* 0 0 256000 0 LISTEN
129.70.161.63.35473 35.232.178.214.443 43008 0 257664 0 ESTABLISHED
*.22 *.* 0 0 256000 0 LISTEN
*.111 *.* 0 0 256000 0 LISTEN
127.0.0.1.5999 *.* 0 0 256000 0 LISTEN
127.0.0.1.6010 *.* 0 0 256000 0 LISTEN
127.0.0.1.39290 *.* 0 0 256000 0 LISTEN
127.0.0.1.50810 127.0.0.1.39290 261760 0 270336 0 ESTABLISHED
127.0.0.1.39290 127.0.0.1.50810 269936 0 261760 0 ESTABLISHED
=== pfiles sockets of the wedged process:
=== pfiles lines mentioning 65381:
# Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3657439566/rsc-gotip-solaris-amd64-0.stdout".
# Error running run: unable to execute /bin/sh: rpc error: code = Unknown desc = command execution failed: exit status 1The connection is completely gone from netstat, yet the goroutine is still parked in netpoll waiting to write. Let me test the specific case: peer destroyed by RST while the writer is armed for POLLOUT.
Ran tools (1 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/rst.go <<'EO...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/rst.go <<'EOF'
// Does a blocked Write wake up when the peer's socket is destroyed?
// Variants: peer closes normally (FIN) vs. with SO_LINGER 0 (RST), and with or
// without a concurrent reader on the server side (net/http has one, but it goes
// away as soon as it sees the FIN/EOF).
package main
import (
"flag"
"fmt"
"io"
"net"
"os"
"time"
)
var wait = flag.Duration("wait", 15*time.Second, "how long to wait for the server side")
func one(rst, reader bool) bool {
ln, err := net.Listen("tcp", "127.0.0.1:0")
if err != nil {
panic(err)
}
defer ln.Close()
readRes := make(chan string, 1)
writeRes := make(chan string, 1)
go func() {
c, err := ln.Accept()
if err != nil {
panic(err)
}
if reader {
go func() {
var b [1]byte
n, err := c.Read(b[:])
readRes <- fmt.Sprintf("n=%d err=%v", n, err)
}()
}
start := time.Now()
n, err := c.Write(make([]byte, 8<<20))
writeRes <- fmt.Sprintf("n=%d err=%v after %v", n, err, time.Since(start).Round(time.Millisecond))
}()
c, err := net.Dial("tcp", ln.Addr().String())
if err != nil {
panic(err)
}
if _, err := io.ReadFull(c, make([]byte, 4096)); err != nil {
panic(err)
}
time.Sleep(200 * time.Millisecond) // server is now blocked in Write
if reader {
// Let the server's reader see the close and go away first, the way
// net/http's connReader.backgroundRead does.
select {
case s := <-readRes:
fmt.Printf(" reader (before close): %s\n", s)
default:
}
}
if rst {
c.(*net.TCPConn).SetLinger(0) // close sends RST
}
c.Close()
select {
case s := <-writeRes:
fmt.Printf(" Write %s\n", s)
if reader {
select {
case s := <-readRes:
fmt.Printf(" Read %s\n", s)
case <-time.After(time.Second):
fmt.Printf(" Read (still blocked)\n")
}
}
return true
case <-time.After(*wait):
fmt.Printf(" Write STUCK after %v\n", *wait)
if reader {
select {
case s := <-readRes:
fmt.Printf(" Read %s <- read side DID see it\n", s)
default:
fmt.Printf(" Read (still blocked)\n")
}
}
return false
}
}
func main() {
flag.Parse()
bad := 0
for _, rst := range []bool{false, true} {
for _, reader := range []bool{false, true} {
fmt.Printf("== close=%s serverReader=%v\n", map[bool]string{false: "FIN", true: "RST"}[rst], reader)
for i := 0; i < 3; i++ {
if !one(rst, reader) {
bad++
}
}
}
}
if bad > 0 {
fmt.Printf("FAIL: %d stuck\n", bad)
os.Exit(1)
}
fmt.Println("PASS")
}
EOF
cd $SP && GOOS=solaris GOARCH=amd64 /home/claude/golang_go_78438/go/bin/go build -o rst.solaris rst.go && G=rsc-gotip-solaris-amd64-0 && gomote put $G rst.solaris rst 2>&1 | tail -1 && gomote run $G /bin/sh -c 'chmod +x rst && ./rst -wait=10s' 2>&1 | tail -40Test RST vs FIN wakeup on solaris
600000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2802140079/rsc-gotip-solaris-amd64-0.stdout"...
== close=FIN serverReader=false
Write n=294912 err=write tcp 127.0.0.1:40139->127.0.0.1:55141: write: broken pipe after 203ms
Write n=294912 err=write tcp 127.0.0.1:35151->127.0.0.1:33609: write: broken pipe after 202ms
Write n=270336 err=write tcp 127.0.0.1:40190->127.0.0.1:59029: write: broken pipe after 217ms
== close=FIN serverReader=true
Write n=294912 err=write tcp 127.0.0.1:55583->127.0.0.1:55535: write: broken pipe after 218ms
Read n=0 err=EOF
Write n=294912 err=write tcp 127.0.0.1:44310->127.0.0.1:44796: write: broken pipe after 205ms
Read n=0 err=read tcp 127.0.0.1:44310->127.0.0.1:44796: read: connection reset by peer
Write n=270336 err=write tcp 127.0.0.1:63228->127.0.0.1:43961: write: broken pipe after 204ms
Read n=0 err=EOF
== close=RST serverReader=false
Write n=294912 err=write tcp 127.0.0.1:33417->127.0.0.1:57490: write: broken pipe after 204ms
Write n=294912 err=write tcp 127.0.0.1:57878->127.0.0.1:51899: write: broken pipe after 204ms
Write n=294912 err=write tcp 127.0.0.1:64632->127.0.0.1:39806: write: broken pipe after 203ms
== close=RST serverReader=true
Write n=294912 err=write tcp 127.0.0.1:60227->127.0.0.1:41425: write: broken pipe after 204ms
Read n=0 err=read tcp 127.0.0.1:60227->127.0.0.1:41425: read: connection reset by peer
Write n=294912 err=write tcp 127.0.0.1:65235->127.0.0.1:63957: write: broken pipe after 202ms
Read n=0 err=read tcp 127.0.0.1:65235->127.0.0.1:63957: read: connection reset by peer
Write n=294912 err=write tcp 127.0.0.1:35808->127.0.0.1:33777: write: broken pipe after 206ms
Read n=0 err=read tcp 127.0.0.1:35808->127.0.0.1:33777: read: connection reset by peer
PASS
# Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2802140079/rsc-gotip-solaris-amd64-0.stdout".
Shell cwd was reset to /home/claude/golang_go_78438All four RST/FIN variants wake up correctly. So it's dynamic — the hang needs the repeated write/wake cycles. Let me A/B test the Solaris netpoll association bookkeeping.
Ran tools (3 Bash)
Bash cd /home/claude/golang_go_78438/go && sed -n '160,182p' src/runtime/netpoll_solaris.go
cd /home/claude/golang_go_78438/go && sed -n '160,182p' src/runtime/netpoll_solaris.go
Show netpollupdate again
// Updates the association with a new set of interested events. After
// this call, port_getn will return one and only one event for that
// particular descriptor, so this function needs to be called again.
func netpollupdate(pd *pollDesc, set, clear uint32) {
if pd.info().closing() {
return
}
old := pd.user
events := (old & ^clear) | set
if old == events {
return
}
tp := taggedPointerPack(unsafe.Pointer(pd), pd.fdseq.Load())
if events != 0 && port_associate(portfd, _PORT_SOURCE_FD, pd.fd, events, uintptr(tp)) != 0 {
print("runtime: port_associate failed (errno=", errno(), ")\n")
throw("runtime: netpollupdate failed")
}
pd.user = events
}Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go
cp src/runtime/netpoll_solaris.go $SP/netpoll_solaris.go.orig
python3 - <<'EOF'
p='src/runtime/netpoll_solaris.go'
s=open(p).read()
old=''' old := pd.user
events := (old & ^clear) | set
if old == events {
return
}
'''
new=''' old := pd.user
events := (old & ^clear) | set
if old == events && set == 0 {
return
}
'''
assert old in s
open(p,'w').write(s.replace(old,new))
EOF
cat > $SP/stress.sh <<'EOF'
#!/bin/sh
# $1 = binary, $2 = tag
bin=$1; tag=$2
rm -f t.$tag.*.out
i=0
while [ $i -lt 8 ]; do
./$bin -test.run '^TestTransportGzip$' -test.count=400 -test.timeout=3m > t.$tag.$i.out 2>&1 &
i=`expr $i + 1`
done
wait
wedged=0
i=0
while [ $i -lt 8 ]; do
if grep -q "test timed out" t.$tag.$i.out; then wedged=`expr $wedged + 1`; fi
i=`expr $i + 1`
done
echo "RESULT $tag: $wedged of 8 wedged"
EOF
GOOS=solaris GOARCH=amd64 /home/claude/golang_go_78438/go/bin/go test -c -o $SP/nethttp.assoc.test net/http 2>&1 | tail -3
cp $SP/netpoll_solaris.go.orig src/runtime/netpoll_solaris.go
git diff --stat src/runtime/netpoll_solaris.go; ls -l $SP/nethttp.assoc.testBuild patched test binary, restore source
300000
-rwxrwxr-x 1 claude claude 17900635 Aug 26 12:12 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/nethttp.assoc.test
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0 gomote put $G $SP/stress.sh stress.sh 2>&1|tail -1 gomote put $G $SP/nethttp.assoc.test nethttp.assoc.test 2>&1|tail -1 gomote run $G /bin/sh -c 'chmod +x stress.sh nethttp.assoc.test; ./stress.sh nethttp.test base; ./stress.sh nethttp.assoc.test assoc' 2>&1 | tail -10
A/B stress: baseline vs always-associate
900000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2847890416/rsc-gotip-solaris-amd64-0.stdout"... RESULT base: 2 of 8 wedged RESULT assoc: 3 of 8 wedged # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2847890416/rsc-gotip-solaris-amd64-0.stdout".
Always-associating doesn't fix it, so the association bookkeeping isn't the cause. Let me see if the standalone model hangs under load.
Ran tools (1 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0 gomote run $G /bin/sh -c 'rm -f r.*.out; i=0; while [ $i -lt 8 ]; do ./solrepro -runs=150 -wait=20s > r.$i.out 2>&1 & i=`expr $i + 1`; done; wait; grep -c STUCK r.*.out; echo ---; tail -1 r.0.out' 2>&1 | tail -15
Run standalone model under load
900000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3704044427/rsc-gotip-solaris-amd64-0.stdout"... r.0.out:0 r.1.out:0 r.2.out:0 r.3.out:0 r.4.out:0 r.5.out:0 r.6.out:0 r.7.out:0 --- PASS # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3704044427/rsc-gotip-solaris-amd64-0.stdout".
Let me instrument internal/poll to find out what the socket would actually do if we retried the write.
Ran tools (3 Bash)
Bash sed -n '380,405p' src/internal/poll/fd_unix.go; echo "=== fd_poll_runtime wait ==="; sed -n '75,100p' src/internal/poll/...
sed -n '380,405p' src/internal/poll/fd_unix.go; echo "=== fd_poll_runtime wait ==="; sed -n '75,100p' src/internal/poll/fd_poll_runtime.go
Show FD.Write and pollDesc.wait
if n > 0 {
if n > max-nn {
// This can reportedly happen when using
// some VPN software. Issue #61060.
// If we don't check this we will panic
// with slice bounds out of range.
// Use a more informative panic.
panic("invalid return from write: got " + strconv.Itoa(n) + " from a write of " + strconv.Itoa(max-nn))
}
nn += n
}
if nn == len(p) {
return nn, err
}
if err == syscall.EAGAIN && fd.pd.pollable() {
if err = fd.pd.waitWrite(fd.isFile); err == nil {
continue
}
}
if err != nil {
return nn, err
}
if n == 0 {
return nn, io.ErrUnexpectedEOF
}
}
=== fd_poll_runtime wait ===
func (pd *pollDesc) prepareWrite(isFile bool) error {
return pd.prepare('w', isFile)
}
func (pd *pollDesc) wait(mode int, isFile bool) error {
if pd.runtimeCtx == 0 {
return errors.New("waiting for unsupported file type")
}
res := runtime_pollWait(pd.runtimeCtx, mode)
return convertErr(res, isFile)
}
func (pd *pollDesc) waitRead(isFile bool) error {
return pd.wait('r', isFile)
}
func (pd *pollDesc) waitWrite(isFile bool) error {
return pd.wait('w', isFile)
}
func (pd *pollDesc) waitCanceled(mode int) {
if pd.runtimeCtx == 0 {
return
}
runtime_pollWaitCanceled(pd.runtimeCtx, mode)Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go
cp src/internal/poll/fd_unix.go $SP/fd_unix.go.orig
cp src/internal/poll/fd_poll_runtime.go $SP/fd_poll_runtime.go.orig
python3 - <<'PY'
p='src/internal/poll/fd_poll_runtime.go'
s=open(p).read()
anchor='''func (pd *pollDesc) waitWrite(isFile bool) error {
return pd.wait('w', isFile)
}
'''
add=anchor+'''
// debugStuckWriteWait bounds how long waitWriteDebug blocks waiting for a
// netpoll write-readiness event before giving up and letting the caller retry
// the write syscall directly. Debugging aid for go.dev/issue/78438.
const debugStuckWriteWait = 15 * time.Second
// waitWriteDebug is waitWrite with a watchdog. It reports whether the wait
// timed out without netpoll ever reporting the fd writable.
func (pd *pollDesc) waitWriteDebug(isFile bool) (error, bool) {
if pd.runtimeCtx == 0 {
return pd.wait('w', isFile), false
}
runtime_pollSetDeadline(pd.runtimeCtx, int64(debugStuckWriteWait), 'w')
err := pd.wait('w', isFile)
runtime_pollSetDeadline(pd.runtimeCtx, 0, 'w')
if err == ErrDeadlineExceeded {
return nil, true
}
return err, false
}
'''
assert anchor in s
s=s.replace(anchor,add)
open(p,'w').write(s)
p='src/internal/poll/fd_unix.go'
s=open(p).read()
old=''' if err == syscall.EAGAIN && fd.pd.pollable() {
if err = fd.pd.waitWrite(fd.isFile); err == nil {
continue
}
}
if err != nil {
return nn, err
}
if n == 0 {
return nn, io.ErrUnexpectedEOF
}
}'''
new=''' if err == syscall.EAGAIN && fd.pd.pollable() {
var stuck bool
if err, stuck = fd.pd.waitWriteDebug(fd.isFile); stuck {
println("STUCKWRITE: fd", fd.Sysfd, "no write event for 15s; retrying write of", len(p)-nn, "bytes")
n2, err2 := ignoringEINTRIO(syscall.Write, fd.Sysfd, p[nn:max])
es := "<nil>"
if err2 != nil {
es = err2.Error()
}
println("STUCKWRITE: fd", fd.Sysfd, "retry returned n =", n2, "err =", es)
}
if err == nil {
continue
}
}
if err != nil {
return nn, err
}
if n == 0 {
return nn, io.ErrUnexpectedEOF
}
}'''
assert old in s
s=s.replace(old,new)
open(p,'w').write(s)
PY
grep -n '"time"' src/internal/poll/fd_poll_runtime.go | head -2
GOOS=solaris GOARCH=amd64 /home/claude/golang_go_78438/go/bin/go test -c -o $SP/nethttp.dbg.test net/http 2>&1 | tail -5
cp $SP/fd_unix.go.orig src/internal/poll/fd_unix.go; cp $SP/fd_poll_runtime.go.orig src/internal/poll/fd_poll_runtime.go
git status --porcelain src/internal/poll | head; ls -l $SP/nethttp.dbg.testBuild instrumented test binary
300000
13: "time" -rwxrwxr-x 1 claude claude 17902252 Aug 26 12:21 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/nethttp.dbg.test
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0 gomote put $G $SP/nethttp.dbg.test nethttp.dbg.test 2>&1|tail -1 gomote run $G /bin/sh -c 'chmod +x nethttp.dbg.test; ./stress.sh nethttp.dbg.test dbg; echo "=== STUCKWRITE lines:"; grep STUCKWRITE t.dbg.*.out | head -30' 2>&1 | tail -30
Run instrumented stress on solaris
900000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote1604583680/rsc-gotip-solaris-amd64-0.stdout"... RESULT dbg: 0 of 8 wedged === STUCKWRITE lines: t.dbg.2.out:STUCKWRITE: fd 8 no write event for 15s; retrying write of 479501 bytes t.dbg.2.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe t.dbg.4.out:STUCKWRITE: fd 8 no write event for 15s; retrying write of 479501 bytes t.dbg.4.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe t.dbg.5.out:STUCKWRITE: fd 8 no write event for 15s; retrying write of 577805 bytes t.dbg.5.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote1604583680/rsc-gotip-solaris-amd64-0.stdout".
Now let me settle whether the kernel owes us an event or Go failed to ask for one.
Ran tools (5 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go grep -n "user uint32" src/runtime/netpoll.go; grep -n "^ fdseq" src/runtime/netpoll.go
Locate pollDesc fields
79: fdseq atomic.Uintptr // protects against stale pollDesc 304: fdseq := pd.fdseq.Load() 305: fdseq = (fdseq + 1) & (1<<tagBits - 1)
Bash sed -n '60,110p' src/runtime/netpoll.go
sed -n '60,110p' src/runtime/netpoll.go
Show pollDesc struct
// G pointer - the goroutine is blocked on the semaphore;
// io notification or timeout/close changes the state to pdReady or pdNil respectively
// and unparks the goroutine.
// pdNil - none of the above.
const (
pdNil uintptr = 0
pdReady uintptr = 1
pdWait uintptr = 2
)
const pollBlockSize = 4 * 1024
// Network poller descriptor.
//
// No heap pointers.
type pollDesc struct {
_ sys.NotInHeap
link *pollDesc // in pollcache, protected by pollcache.lock
fd uintptr // constant for pollDesc usage lifetime
fdseq atomic.Uintptr // protects against stale pollDesc
// atomicInfo holds bits from closing, rd, and wd,
// which are only ever written while holding the lock,
// summarized for use by netpollcheckerr,
// which cannot acquire the lock.
// After writing these fields under lock in a way that
// might change the summary, code must call publishInfo
// before releasing the lock.
// Code that changes fields and then calls netpollunblock
// (while still holding the lock) must call publishInfo
// before calling netpollunblock, because publishInfo is what
// stops netpollblock from blocking anew
// (by changing the result of netpollcheckerr).
// atomicInfo also holds the eventErr bit,
// recording whether a poll event on the fd got an error;
// atomicInfo is the only source of truth for that bit.
atomicInfo atomic.Uint32 // atomic pollInfo
// rg, wg are accessed atomically and hold g pointers.
// (Using atomic.Uintptr here is similar to using guintptr elsewhere.)
rg atomic.Uintptr // pdReady, pdWait, G waiting for read or pdNil
wg atomic.Uintptr // pdReady, pdWait, G waiting for write or pdNil
lock mutex // protects the following fields
closing bool
rrun bool // whether rt is running
wrun bool // whether wt is running
user uint32 // user settable cookie
rseq uintptr // protects from stale read timers
rt timer // read deadline timer
rd int64 // read deadline (a nanotime in the future, -1 when expired)Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go
cp src/runtime/netpoll.go $SP/netpoll.go.orig
python3 - <<'PY'
# 1. runtime/netpoll.go: add debug counters + accessor
p='src/runtime/netpoll.go'
s=open(p).read()
old=''' user uint32 // user settable cookie
'''
new=''' user uint32 // user settable cookie
dbgAssoc uint32 // debug: successful port_associate calls (solaris)
dbgEv uint32 // debug: events delivered for this pd (solaris)
dbgLast uint32 // debug: portev_events of the last delivered event
'''
assert old in s
s=s.replace(old,new,1)
s=s+'''
// poll_runtime_pollDebug reports netpoll bookkeeping for pd, for debugging
// go.dev/issue/78438.
//
//go:linkname poll_runtime_pollDebug internal/poll.runtime_pollDebug
func poll_runtime_pollDebug(ctx uintptr) (user, assoc, ev, last uint32) {
pd := (*pollDesc)(unsafe.Pointer(ctx))
lock(&pd.lock)
user, assoc, ev, last = pd.user, pd.dbgAssoc, pd.dbgEv, pd.dbgLast
unlock(&pd.lock)
return
}
'''
open(p,'w').write(s)
# 2. runtime/netpoll_solaris.go: bump counters
p='src/runtime/netpoll_solaris.go'
s=open(p).read()
old=''' if events != 0 && port_associate(portfd, _PORT_SOURCE_FD, pd.fd, events, uintptr(tp)) != 0 {
print("runtime: port_associate failed (errno=", errno(), ")\\n")
throw("runtime: netpollupdate failed")
}
pd.user = events'''
new=''' if events != 0 {
if port_associate(portfd, _PORT_SOURCE_FD, pd.fd, events, uintptr(tp)) != 0 {
print("runtime: port_associate failed (errno=", errno(), ")\\n")
throw("runtime: netpollupdate failed")
}
pd.dbgAssoc++
}
pd.user = events'''
assert old in s
s=s.replace(old,new,1)
old2=''' if pd.fdseq.Load() != tp.tag() {
continue
}
'''
new2=''' if pd.fdseq.Load() != tp.tag() {
continue
}
lock(&pd.lock)
pd.dbgEv++
pd.dbgLast = ev.portev_events
unlock(&pd.lock)
'''
assert old2 in s
s=s.replace(old2,new2,1)
open(p,'w').write(s)
# 3. internal/poll: declare + report
p='src/internal/poll/fd_poll_runtime.go'
s=open(p).read()
old='''func runtime_pollWait(ctx uintptr, mode int) int'''
new='''func runtime_pollWait(ctx uintptr, mode int) int
func runtime_pollDebug(ctx uintptr) (user, assoc, ev, last uint32)'''
assert old in s
s=s.replace(old,new,1)
anchor='''func (pd *pollDesc) waitWrite(isFile bool) error {
return pd.wait('w', isFile)
}
'''
add=anchor+'''
const debugStuckWriteWait = 15 * time.Second
func (pd *pollDesc) waitWriteDebug(isFile bool) (error, bool) {
if pd.runtimeCtx == 0 {
return pd.wait('w', isFile), false
}
runtime_pollSetDeadline(pd.runtimeCtx, int64(debugStuckWriteWait), 'w')
err := pd.wait('w', isFile)
runtime_pollSetDeadline(pd.runtimeCtx, 0, 'w')
if err == ErrDeadlineExceeded {
return nil, true
}
return err, false
}
func (pd *pollDesc) debugState() (user, assoc, ev, last uint32) {
return runtime_pollDebug(pd.runtimeCtx)
}
'''
assert anchor in s
s=s.replace(anchor,add,1)
open(p,'w').write(s)
p='src/internal/poll/fd_unix.go'
s=open(p).read()
old=''' if err == syscall.EAGAIN && fd.pd.pollable() {
if err = fd.pd.waitWrite(fd.isFile); err == nil {
continue
}
}
if err != nil {
return nn, err
}
if n == 0 {
return nn, io.ErrUnexpectedEOF
}
}'''
new=''' if err == syscall.EAGAIN && fd.pd.pollable() {
var stuck bool
if err, stuck = fd.pd.waitWriteDebug(fd.isFile); stuck {
user, assoc, ev, last := fd.pd.debugState()
println("STUCKWRITE: fd", fd.Sysfd, "no write event for 15s; remaining", len(p)-nn,
"| pd.user =", user, "assocs =", assoc, "events =", ev, "lastEvents =", last)
n2, err2 := ignoringEINTRIO(syscall.Write, fd.Sysfd, p[nn:max])
es := "<nil>"
if err2 != nil {
es = err2.Error()
}
println("STUCKWRITE: fd", fd.Sysfd, "retry returned n =", n2, "err =", es)
}
if err == nil {
continue
}
}
if err != nil {
return nn, err
}
if n == 0 {
return nn, io.ErrUnexpectedEOF
}
}'''
assert old in s
s=s.replace(old,new,1)
open(p,'w').write(s)
PY
GOOS=solaris GOARCH=amd64 /home/claude/golang_go_78438/go/bin/go test -c -o $SP/nethttp.dbg2.test net/http 2>&1 | tail -5
cp $SP/netpoll.go.orig src/runtime/netpoll.go
cp $SP/netpoll_solaris.go.orig src/runtime/netpoll_solaris.go
cp $SP/fd_unix.go.orig src/internal/poll/fd_unix.go
cp $SP/fd_poll_runtime.go.orig src/internal/poll/fd_poll_runtime.go
git status --porcelain src/runtime src/internal/poll | head; ls -l $SP/nethttp.dbg2.testBuild instrumented binary with netpoll bookkeeping
300000
Exit code 2 # runtime src/runtime/netpoll_solaris.go:308:16: cannot use ev.portev_events (variable of type int32) as uint32 value in assignment ls: cannot access '/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/nethttp.dbg2.test': No such file or directory
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go
python3 - <<'PY'
p='src/runtime/netpoll.go'
s=open(p).read()
old=''' user uint32 // user settable cookie
'''
new=''' user uint32 // user settable cookie
dbgAssoc uint32 // debug: successful port_associate calls (solaris)
dbgEv uint32 // debug: events delivered for this pd (solaris)
dbgLast uint32 // debug: portev_events of the last delivered event
'''
assert old in s
s=s.replace(old,new,1)
s=s+'''
//go:linkname poll_runtime_pollDebug internal/poll.runtime_pollDebug
func poll_runtime_pollDebug(ctx uintptr) (user, assoc, ev, last uint32) {
pd := (*pollDesc)(unsafe.Pointer(ctx))
lock(&pd.lock)
user, assoc, ev, last = pd.user, pd.dbgAssoc, pd.dbgEv, pd.dbgLast
unlock(&pd.lock)
return
}
'''
open(p,'w').write(s)
p='src/runtime/netpoll_solaris.go'
s=open(p).read()
old=''' if events != 0 && port_associate(portfd, _PORT_SOURCE_FD, pd.fd, events, uintptr(tp)) != 0 {
print("runtime: port_associate failed (errno=", errno(), ")\\n")
throw("runtime: netpollupdate failed")
}
pd.user = events'''
new=''' if events != 0 {
if port_associate(portfd, _PORT_SOURCE_FD, pd.fd, events, uintptr(tp)) != 0 {
print("runtime: port_associate failed (errno=", errno(), ")\\n")
throw("runtime: netpollupdate failed")
}
pd.dbgAssoc++
}
pd.user = events'''
assert old in s
s=s.replace(old,new,1)
old2=''' if pd.fdseq.Load() != tp.tag() {
continue
}
'''
new2=''' if pd.fdseq.Load() != tp.tag() {
continue
}
lock(&pd.lock)
pd.dbgEv++
pd.dbgLast = uint32(ev.portev_events)
unlock(&pd.lock)
'''
assert old2 in s
s=s.replace(old2,new2,1)
open(p,'w').write(s)
p='src/internal/poll/fd_poll_runtime.go'
s=open(p).read()
old='''func runtime_pollWait(ctx uintptr, mode int) int'''
new='''func runtime_pollWait(ctx uintptr, mode int) int
func runtime_pollDebug(ctx uintptr) (user, assoc, ev, last uint32)'''
assert old in s
s=s.replace(old,new,1)
anchor='''func (pd *pollDesc) waitWrite(isFile bool) error {
return pd.wait('w', isFile)
}
'''
add=anchor+'''
const debugStuckWriteWait = 15 * time.Second
func (pd *pollDesc) waitWriteDebug(isFile bool) (error, bool) {
if pd.runtimeCtx == 0 {
return pd.wait('w', isFile), false
}
runtime_pollSetDeadline(pd.runtimeCtx, int64(debugStuckWriteWait), 'w')
err := pd.wait('w', isFile)
runtime_pollSetDeadline(pd.runtimeCtx, 0, 'w')
if err == ErrDeadlineExceeded {
return nil, true
}
return err, false
}
func (pd *pollDesc) debugState() (user, assoc, ev, last uint32) {
return runtime_pollDebug(pd.runtimeCtx)
}
'''
assert anchor in s
s=s.replace(anchor,add,1)
open(p,'w').write(s)
p='src/internal/poll/fd_unix.go'
s=open(p).read()
old=''' if err == syscall.EAGAIN && fd.pd.pollable() {
if err = fd.pd.waitWrite(fd.isFile); err == nil {
continue
}
}
if err != nil {
return nn, err
}
if n == 0 {
return nn, io.ErrUnexpectedEOF
}
}'''
new=''' if err == syscall.EAGAIN && fd.pd.pollable() {
var stuck bool
if err, stuck = fd.pd.waitWriteDebug(fd.isFile); stuck {
user, assoc, ev, last := fd.pd.debugState()
println("STUCKWRITE: fd", fd.Sysfd, "no write event for 15s; remaining", len(p)-nn,
"| pd.user =", user, "assocs =", assoc, "events =", ev, "lastEvents =", last)
n2, err2 := ignoringEINTRIO(syscall.Write, fd.Sysfd, p[nn:max])
es := "<nil>"
if err2 != nil {
es = err2.Error()
}
println("STUCKWRITE: fd", fd.Sysfd, "retry returned n =", n2, "err =", es)
}
if err == nil {
continue
}
}
if err != nil {
return nn, err
}
if n == 0 {
return nn, io.ErrUnexpectedEOF
}
}'''
assert old in s
s=s.replace(old,new,1)
open(p,'w').write(s)
PY
GOOS=solaris GOARCH=amd64 /home/claude/golang_go_78438/go/bin/go test -c -o $SP/nethttp.dbg2.test net/http 2>&1 | tail -5
cp $SP/netpoll.go.orig src/runtime/netpoll.go
cp $SP/netpoll_solaris.go.orig src/runtime/netpoll_solaris.go
cp $SP/fd_unix.go.orig src/internal/poll/fd_unix.go
cp $SP/fd_poll_runtime.go.orig src/internal/poll/fd_poll_runtime.go
git status --porcelain src/runtime src/internal/poll | head; ls -l $SP/nethttp.dbg2.testRebuild instrumented binary
300000
-rwxrwxr-x 1 claude claude 17903156 Aug 26 12:23 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/nethttp.dbg2.test
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0 gomote put $G $SP/nethttp.dbg2.test nethttp.dbg2.test 2>&1|tail -1 gomote run $G /bin/sh -c 'chmod +x nethttp.dbg2.test; ./stress.sh nethttp.dbg2.test dbg2; echo "=== STUCKWRITE:"; grep STUCKWRITE t.dbg2.*.out | head -20' 2>&1 | tail -25
Run instrumented stress with netpoll state
900000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2941693178/rsc-gotip-solaris-amd64-0.stdout"... RESULT dbg2: 0 of 8 wedged === STUCKWRITE: t.dbg2.2.out:STUCKWRITE: fd 8 no write event for 15s; remaining 454925 | pd.user = 4 assocs = 392 events = 337 lastEvents = 1 t.dbg2.2.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe t.dbg2.2.out:STUCKWRITE: fd 8 no write event for 15s; remaining 430349 | pd.user = 4 assocs = 3888 events = 3393 lastEvents = 1 t.dbg2.2.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe t.dbg2.2.out:STUCKWRITE: fd 8 no write event for 15s; remaining 454925 | pd.user = 4 assocs = 4175 events = 3640 lastEvents = 1 t.dbg2.2.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe t.dbg2.4.out:STUCKWRITE: fd 8 no write event for 15s; remaining 479501 | pd.user = 4 assocs = 4837 events = 4249 lastEvents = 1 t.dbg2.4.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe t.dbg2.5.out:STUCKWRITE: fd 8 no write event for 15s; remaining 405773 | pd.user = 4 assocs = 2335 events = 2030 lastEvents = 1 t.dbg2.5.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe t.dbg2.7.out:STUCKWRITE: fd 8 no write event for 15s; remaining 405773 | pd.user = 4 assocs = 993 events = 864 lastEvents = 1 t.dbg2.7.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe t.dbg2.7.out:STUCKWRITE: fd 8 no write event for 15s; remaining 528653 | pd.user = 4 assocs = 3585 events = 3140 lastEvents = 1 t.dbg2.7.out:STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2941693178/rsc-gotip-solaris-amd64-0.stdout".
Now the picture is sharp: pd.user = 4 (POLLOUT armed), last delivered event was 1 (POLLIN only), and a direct write returns EPIPE instantly. Let me confirm this at the kernel level with a C program using event ports directly.
Ran tools (4 Bash)
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'which cc gcc; ls /usr/bin/*cc* /usr/gcc 2>/dev/null | head' 2>&1 ...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'which cc gcc; ls /usr/bin/*cc* /usr/gcc 2>/dev/null | head' 2>&1 | tail -10
Check for a C compiler on the buildlet
/usr/bin/ccmake /usr/bin/cd-iccdump /usr/bin/gcc /usr/bin/gccgo /usr/bin/jpgicc /usr/bin/linkicc /usr/bin/orcc /usr/bin/psicc /usr/bin/rpcclient # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2700120688/rsc-gotip-solaris-amd64-0.stdout".
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/portbug.c <<...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/portbug.c <<'EOF'
/*
* Does a Solaris event port ever report a socket as writable (or in error)
* once the peer is gone and the pending socket error has already been
* consumed by a read?
*
* This is what Go's netpoll does on solaris (runtime/netpoll_solaris.go):
* an fd is port_associate'd with the event set the blocked goroutines want.
* When an event is delivered the fd is dissociated, so the poller
* re-associates with the events that were not part of that notification.
* A writer blocked with POLLOUT armed is therefore woken only if the port
* reports POLLOUT/POLLERR/POLLHUP for it. See go.dev/issue/78438.
*
* Build: gcc -o portbug portbug.c -lsocket -lnsl
*/
#include <errno.h>
#include <fcntl.h>
#include <netinet/in.h>
#include <netinet/tcp.h>
#include <poll.h>
#include <port.h>
#include <stdio.h>
#include <string.h>
#include <stdlib.h>
#include <sys/socket.h>
#include <unistd.h>
static void die(const char *m) { perror(m); exit(2); }
/* consumeError: read() the dead connection first, the way net/http's
* background read goroutine does, before arming for write. */
static int run(int rst, int consumeError)
{
int lfd, cfd, sfd, port, rc, nassoc = 0;
struct sockaddr_in sa;
socklen_t salen;
char *buf;
ssize_t n;
long total = 0;
port_event_t pe;
timespec_t ts;
struct linger lg;
struct pollfd pfd;
printf("== peer close = %s, error consumed by read() first = %s\n",
rst ? "RST (SO_LINGER 0)" : "FIN", consumeError ? "yes" : "no");
lfd = socket(AF_INET, SOCK_STREAM, 0);
if (lfd < 0) die("socket");
memset(&sa, 0, sizeof sa);
sa.sin_family = AF_INET;
sa.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
if (bind(lfd, (struct sockaddr *)&sa, sizeof sa) < 0) die("bind");
if (listen(lfd, 1) < 0) die("listen");
salen = sizeof sa;
if (getsockname(lfd, (struct sockaddr *)&sa, &salen) < 0) die("getsockname");
cfd = socket(AF_INET, SOCK_STREAM, 0);
if (cfd < 0) die("socket");
if (connect(cfd, (struct sockaddr *)&sa, salen) < 0) die("connect");
sfd = accept(lfd, NULL, NULL);
if (sfd < 0) die("accept");
close(lfd);
if (fcntl(sfd, F_SETFL, O_NONBLOCK) < 0) die("fcntl");
/* Fill the send buffer: write until EAGAIN, like a blocked handler. */
buf = malloc(1 << 20);
memset(buf, 'x', 1 << 20);
for (;;) {
n = write(sfd, buf, 1 << 20);
if (n < 0) {
if (errno == EAGAIN) break;
die("write(fill)");
}
total += n;
}
printf(" filled send buffer with %ld bytes, then EAGAIN\n", total);
/* Kill the peer. */
if (rst) {
lg.l_onoff = 1;
lg.l_linger = 0;
if (setsockopt(cfd, SOL_SOCKET, SO_LINGER, &lg, sizeof lg) < 0)
die("SO_LINGER");
}
close(cfd);
sleep(1);
if (consumeError) {
char b[1];
n = read(sfd, b, 1);
printf(" read() after peer close: n=%zd errno=%d (%s)\n",
n, n < 0 ? errno : 0, n < 0 ? strerror(errno) : "-");
}
/* What does plain poll(2) say? (does not consume anything) */
pfd.fd = sfd;
pfd.events = POLLOUT;
pfd.revents = 0;
rc = poll(&pfd, 1, 0);
printf(" poll(POLLOUT) -> rc=%d revents=0x%x\n", rc, pfd.revents);
/* Now do exactly what Go's netpoll does. */
port = port_create();
if (port < 0) die("port_create");
if (port_associate(port, PORT_SOURCE_FD, sfd, 0, NULL) < 0)
die("port_associate(0)"); /* netpollopen */
if (port_associate(port, PORT_SOURCE_FD, sfd, POLLOUT, NULL) < 0)
die("port_associate(POLLOUT)"); /* netpollarm('w') */
nassoc++;
ts.tv_sec = 5;
ts.tv_nsec = 0;
memset(&pe, 0, sizeof pe);
rc = port_get(port, &pe, &ts);
if (rc < 0 && errno == ETIME)
printf(" port_get(POLLOUT) -> TIMED OUT after 5s <<< no event\n");
else if (rc < 0)
printf(" port_get(POLLOUT) -> errno=%d (%s)\n", errno, strerror(errno));
else
printf(" port_get(POLLOUT) -> events=0x%x\n", pe.portev_events);
n = write(sfd, buf, 1);
printf(" write() right afterwards: n=%zd errno=%d (%s)\n",
n, n < 0 ? errno : 0, n < 0 ? strerror(errno) : "-");
close(port);
close(sfd);
free(buf);
printf("\n");
return 0;
}
int main(void)
{
run(0, 0);
run(0, 1);
run(1, 0);
run(1, 1);
return 0;
}
EOF
G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/portbug.c portbug.c 2>&1|tail -1; gomote run $G /bin/sh -c 'gcc -o portbug portbug.c -lsocket -lnsl && ./portbug' 2>&1 | tail -40Build and run C event-port test on solaris
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote331129744/rsc-gotip-solaris-amd64-0.stdout"... # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote331129744/rsc-gotip-solaris-amd64-0.stdout". # Error running run: unable to execute /bin/sh: rpc error: code = Unknown desc = command execution failed: signal: broken Pipe
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'gcc -o portbug portbug.c -lsocket -lnsl 2>&1 | head -20; echo "bu...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'gcc -o portbug portbug.c -lsocket -lnsl 2>&1 | head -20; echo "build rc=$?"; ls -l portbug 2>&1' 2>&1 | tail -20
Compile C program separately
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2335387971/rsc-gotip-solaris-amd64-0.stdout"... build rc=0 -rwxr-xr-x 1 swarming golang 15896 Aug 26 18:27 portbug # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2335387971/rsc-gotip-solaris-amd64-0.stdout".
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; python3 - <<'PY' ...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; python3 - <<'PY'
p="/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/portbug.c"
s=open(p).read()
s=s.replace('#include <errno.h>','#include <errno.h>\n#include <signal.h>')
s=s.replace('int main(void)\n{\n','int main(void)\n{\n\tsignal(SIGPIPE, SIG_IGN);\n')
open(p,'w').write(s)
PY
G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/portbug.c portbug.c 2>&1|tail -1; gomote run $G /bin/sh -c 'gcc -o portbug portbug.c -lsocket -lnsl && ./portbug' 2>&1 | tail -40Rerun C test with SIGPIPE ignored
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote4130654817/rsc-gotip-solaris-amd64-0.stdout"... == peer close = FIN, error consumed by read() first = no filled send buffer with 270336 bytes, then EAGAIN poll(POLLOUT) -> rc=1 revents=0x4 port_get(POLLOUT) -> events=0x4 write() right afterwards: n=1 errno=0 (-) == peer close = FIN, error consumed by read() first = yes filled send buffer with 270336 bytes, then EAGAIN read() after peer close: n=0 errno=0 (-) poll(POLLOUT) -> rc=1 revents=0x4 port_get(POLLOUT) -> events=0x4 write() right afterwards: n=1 errno=0 (-) == peer close = RST (SO_LINGER 0), error consumed by read() first = no filled send buffer with 270336 bytes, then EAGAIN poll(POLLOUT) -> rc=1 revents=0x4 port_get(POLLOUT) -> events=0x4 write() right afterwards: n=-1 errno=32 (Broken pipe) == peer close = RST (SO_LINGER 0), error consumed by read() first = yes filled send buffer with 270336 bytes, then EAGAIN read() after peer close: n=-1 errno=131 (Connection reset by peer) poll(POLLOUT) -> rc=0 revents=0x0 port_get(POLLOUT) -> TIMED OUT after 5s <<< no event write() right afterwards: n=-1 errno=32 (Broken pipe) # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote4130654817/rsc-gotip-solaris-amd64-0.stdout".
Reproduced with no Go involved. Let me isolate the exact preconditions.
Ran tools (1 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/portbug2.c <...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/portbug2.c <<'EOF'
/*
* solaris: after a socket's pending error has been consumed by read(),
* poll(2) and event ports report the fd as neither writable nor in error,
* even though write() fails immediately with EPIPE. go.dev/issue/78438
*
* Build: gcc -o portbug2 portbug2.c -lsocket -lnsl
*/
#include <errno.h>
#include <fcntl.h>
#include <netinet/in.h>
#include <poll.h>
#include <port.h>
#include <signal.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/socket.h>
#include <unistd.h>
static void die(const char *m) { perror(m); exit(2); }
static void run(int fill, int consume)
{
int lfd, cfd, sfd, port, rc, soerr;
socklen_t slen;
struct sockaddr_in sa;
socklen_t salen;
char *buf;
ssize_t n;
long total = 0;
port_event_t pe;
timespec_t ts;
struct linger lg;
struct pollfd pfd;
printf("== send buffer filled = %-3s error consumed by read() = %s\n",
fill ? "yes" : "no", consume ? "yes" : "no");
lfd = socket(AF_INET, SOCK_STREAM, 0);
memset(&sa, 0, sizeof sa);
sa.sin_family = AF_INET;
sa.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
if (bind(lfd, (struct sockaddr *)&sa, sizeof sa) < 0) die("bind");
if (listen(lfd, 1) < 0) die("listen");
salen = sizeof sa;
getsockname(lfd, (struct sockaddr *)&sa, &salen);
cfd = socket(AF_INET, SOCK_STREAM, 0);
if (connect(cfd, (struct sockaddr *)&sa, salen) < 0) die("connect");
sfd = accept(lfd, NULL, NULL);
close(lfd);
fcntl(sfd, F_SETFL, O_NONBLOCK);
buf = malloc(1 << 20);
memset(buf, 'x', 1 << 20);
if (fill) {
for (;;) {
n = write(sfd, buf, 1 << 20);
if (n < 0) { if (errno == EAGAIN) break; die("write"); }
total += n;
}
printf(" send buffer: %ld bytes queued, now EAGAIN\n", total);
}
lg.l_onoff = 1; lg.l_linger = 0; /* close sends RST */
setsockopt(cfd, SOL_SOCKET, SO_LINGER, &lg, sizeof lg);
close(cfd);
sleep(1);
if (consume) {
char b[1];
n = read(sfd, b, 1);
printf(" read() -> n=%zd errno=%d (%s)\n", n,
n < 0 ? errno : 0, n < 0 ? strerror(errno) : "-");
}
pfd.fd = sfd;
pfd.events = POLLIN | POLLOUT;
pfd.revents = 0;
rc = poll(&pfd, 1, 0);
printf(" poll(IN|OUT) -> rc=%d revents=0x%x%s\n", rc, pfd.revents,
rc == 0 ? " <<< reports nothing" : "");
port = port_create();
port_associate(port, PORT_SOURCE_FD, sfd, 0, NULL);
if (port_associate(port, PORT_SOURCE_FD, sfd, POLLOUT, NULL) < 0)
die("port_associate");
ts.tv_sec = 5; ts.tv_nsec = 0;
memset(&pe, 0, sizeof pe);
rc = port_get(port, &pe, &ts);
if (rc < 0 && errno == ETIME)
printf(" port_get(OUT) -> TIMED OUT after 5s <<< no event, ever\n");
else if (rc < 0)
printf(" port_get(OUT) -> errno=%d (%s)\n", errno, strerror(errno));
else
printf(" port_get(OUT) -> events=0x%x\n", pe.portev_events);
n = write(sfd, buf, 1);
printf(" write() -> n=%zd errno=%d (%s)\n", n,
n < 0 ? errno : 0, n < 0 ? strerror(errno) : "-");
slen = sizeof soerr;
getsockopt(sfd, SOL_SOCKET, SO_ERROR, &soerr, &slen);
printf(" SO_ERROR -> %d (%s)\n", soerr, soerr ? strerror(soerr) : "-");
close(port);
close(sfd);
free(buf);
printf("\n");
}
int main(void)
{
signal(SIGPIPE, SIG_IGN);
run(0, 0);
run(0, 1);
run(1, 0);
run(1, 1);
return 0;
}
EOF
G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/portbug2.c portbug2.c 2>&1|tail -1; gomote run $G /bin/sh -c 'gcc -o portbug2 portbug2.c -lsocket -lnsl && ./portbug2' 2>&1 | tail -45Isolate preconditions in C
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote1675953356/rsc-gotip-solaris-amd64-0.stdout"... == send buffer filled = no error consumed by read() = no poll(IN|OUT) -> rc=1 revents=0x5 port_get(OUT) -> events=0x4 write() -> n=-1 errno=32 (Broken pipe) SO_ERROR -> 131 (Connection reset by peer) == send buffer filled = no error consumed by read() = yes read() -> n=-1 errno=131 (Connection reset by peer) poll(IN|OUT) -> rc=1 revents=0x1 port_get(OUT) -> TIMED OUT after 5s <<< no event, ever write() -> n=-1 errno=32 (Broken pipe) SO_ERROR -> 0 (-) == send buffer filled = yes error consumed by read() = no send buffer: 270336 bytes queued, now EAGAIN poll(IN|OUT) -> rc=1 revents=0x5 port_get(OUT) -> events=0x4 write() -> n=-1 errno=32 (Broken pipe) SO_ERROR -> 131 (Connection reset by peer) == send buffer filled = yes error consumed by read() = yes send buffer: 270336 bytes queued, now EAGAIN read() -> n=-1 errno=131 (Connection reset by peer) poll(IN|OUT) -> rc=1 revents=0x1 port_get(OUT) -> TIMED OUT after 5s <<< no event, ever write() -> n=-1 errno=32 (Broken pipe) SO_ERROR -> 0 (-) # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote1675953356/rsc-gotip-solaris-amd64-0.stdout".
The C repro pins it exactly. Now let me test a Go-side mitigation.
Ran tools (7 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go
cp src/net/http/server.go $SP/server.go.orig
python3 - <<'PY'
p='src/net/http/server.go'
s=open(p).read()
old='''func (cr *connReader) handleReadErrorLocked(_ error) {
if cr.conn == nil {
return
}
cr.conn.cancelCtx()'''
new='''func (cr *connReader) handleReadErrorLocked(err error) {
if cr.conn == nil {
return
}
// A hard read error means the connection is gone, not merely
// half-closed, so abort any response write in flight. On some systems
// (solaris) a write blocked in the poller is never reported as ready
// once a read has consumed the socket's pending error, so the handler
// would otherwise block forever. See go.dev/issue/78438.
if err != nil && err != io.EOF {
if ne, ok := err.(net.Error); !ok || !ne.Timeout() {
cr.conn.rwc.SetWriteDeadline(aLongTimeAgo)
}
}
cr.conn.cancelCtx()'''
assert old in s
open(p,'w').write(s.replace(old,new,1))
PY
GOOS=solaris GOARCH=amd64 /home/claude/golang_go_78438/go/bin/go test -c -o $SP/nethttp.fix.test net/http 2>&1 | tail -5
git diff --stat src/net/http/server.go
cp $SP/server.go.orig src/net/http/server.go
ls -l $SP/nethttp.fix.testBuild candidate net/http mitigation
300000
src/net/http/server.go | 12 +++++++++++- 1 file changed, 11 insertions(+), 1 deletion(-) -rwxrwxr-x 1 claude claude 17901268 Aug 26 12:29 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/nethttp.fix.test
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0 gomote put $G $SP/nethttp.fix.test nethttp.fix.test 2>&1|tail -1 gomote run $G /bin/sh -c 'chmod +x nethttp.fix.test; ./stress.sh nethttp.test base2; ./stress.sh nethttp.fix.test fix' 2>&1 | tail -6
A/B stress: baseline vs mitigation
900000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2505476894/rsc-gotip-solaris-amd64-0.stdout"... RESULT base2: 4 of 8 wedged RESULT fix: 0 of 8 wedged # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2505476894/rsc-gotip-solaris-amd64-0.stdout".
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0 gomote run $G /bin/sh -c './stress.sh nethttp.fix.test fix2; ./stress.sh nethttp.fix.test fix3' 2>&1 | tail -5
Repeat mitigation stress
900000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2248984430/rsc-gotip-solaris-amd64-0.stdout"... RESULT fix2: 0 of 8 wedged RESULT fix3: 0 of 8 wedged # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2248984430/rsc-gotip-solaris-amd64-0.stdout".
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'for b in nethttp.test nethttp.fix.test; do echo "===== $b"; ./$b ...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'for b in nethttp.test nethttp.fix.test; do echo "===== $b"; ./$b -test.timeout=12m > full.$b.out 2>&1; echo "exit=$?"; grep "^--- FAIL" full.$b.out | sort | head -40; grep -c "^--- FAIL" full.$b.out; done' 2>&1 | tail -40
Full net/http suite: baseline vs mitigation
1800000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2099641424/rsc-gotip-solaris-amd64-0.stdout"... ===== nethttp.test exit=1 --- FAIL: TestEmptyDirOpenCWD (0.00s) --- FAIL: TestFileServerDirWithRootFile (0.02s) --- FAIL: TestFileServerMethods (0.00s) --- FAIL: TestFSRedirect (0.02s) --- FAIL: TestNoUnicodeStrings (15.67s) --- FAIL: TestServeContent (0.01s) --- FAIL: TestServeFile (0.01s) --- FAIL: TestServeFile_DotDot (0.00s) --- FAIL: TestServeFileContentType (0.01s) --- FAIL: TestServeFileDirPanicEmptyPath (0.00s) --- FAIL: TestServeFileFromCWD (0.01s) --- FAIL: TestServeFileMimeType (0.01s) --- FAIL: TestServeFileNotModified (0.01s) --- FAIL: TestServeFileRejectsInvalidSuffixLengths (0.03s) --- FAIL: TestServeIndexHtml (0.02s) --- FAIL: TestServeIndexHtmlFS (0.01s) 16 ===== nethttp.fix.test exit=1 --- FAIL: TestEmptyDirOpenCWD (0.00s) --- FAIL: TestFileServerDirWithRootFile (0.02s) --- FAIL: TestFileServerMethods (0.00s) --- FAIL: TestFSRedirect (0.02s) --- FAIL: TestNoUnicodeStrings (8.42s) --- FAIL: TestServeContent (0.00s) --- FAIL: TestServeFile (0.01s) --- FAIL: TestServeFile_DotDot (0.00s) --- FAIL: TestServeFileContentType (0.02s) --- FAIL: TestServeFileDirPanicEmptyPath (0.00s) --- FAIL: TestServeFileFromCWD (0.02s) --- FAIL: TestServeFileMimeType (0.01s) --- FAIL: TestServeFileNotModified (0.01s) --- FAIL: TestServeFileRejectsInvalidSuffixLengths (0.03s) --- FAIL: TestServeIndexHtml (0.02s) --- FAIL: TestServeIndexHtmlFS (0.01s) 16 # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2099641424/rsc-gotip-solaris-amd64-0.stdout".
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go
python3 - <<'PY'
p='src/net/http/server.go'
s=open(p).read()
old='''func (cr *connReader) handleReadErrorLocked(_ error) {
if cr.conn == nil {
return
}
cr.conn.cancelCtx()'''
new='''func (cr *connReader) handleReadErrorLocked(err error) {
if cr.conn == nil {
return
}
// A hard read error means the connection is gone, not merely
// half-closed, so abort any response write in flight. On some systems
// (solaris) a write blocked in the poller is never reported as ready
// once a read has consumed the socket's pending error, so the handler
// would otherwise block forever. See go.dev/issue/78438.
if err != nil && err != io.EOF {
if ne, ok := err.(net.Error); !ok || !ne.Timeout() {
cr.conn.rwc.SetWriteDeadline(aLongTimeAgo)
}
}
cr.conn.cancelCtx()'''
assert old in s
open(p,'w').write(s.replace(old,new,1))
PY
cd src && ../bin/go test -count=1 net/http net/http/httputil net/http/httptest 2>&1 | tail -15Run net/http tests on linux with the patch
900000
ok net/http 11.961s ok net/http/httputil 0.627s ok net/http/httptest 0.029s
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cd /home/claude/golang_go_78438/go && git diff > $SP/proposed-fix.diff; cat $SP/proposed-fix.diff; echo "=== other modified files:"; git status --porcelain | head
Show the candidate diff and repo state
diff --git a/src/net/http/server.go b/src/net/http/server.go
index 513709fa89..9736f5cce6 100644
--- a/src/net/http/server.go
+++ b/src/net/http/server.go
@@ -810,10 +810,20 @@ func (cr *connReader) hitReadLimit() bool { return cr.remain <= 0 }
// down its context.
//
// The caller must hold connReader.mu.
-func (cr *connReader) handleReadErrorLocked(_ error) {
+func (cr *connReader) handleReadErrorLocked(err error) {
if cr.conn == nil {
return
}
+ // A hard read error means the connection is gone, not merely
+ // half-closed, so abort any response write in flight. On some systems
+ // (solaris) a write blocked in the poller is never reported as ready
+ // once a read has consumed the socket's pending error, so the handler
+ // would otherwise block forever. See go.dev/issue/78438.
+ if err != nil && err != io.EOF {
+ if ne, ok := err.(net.Error); !ok || !ne.Timeout() {
+ cr.conn.rwc.SetWriteDeadline(aLongTimeAgo)
+ }
+ }
cr.conn.cancelCtx()
if res := cr.conn.curReq.Load(); res != nil {
res.closeNotify()
=== other modified files:
M src/net/http/server.goBash gomote list 2>&1 | grep solaris; ls /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/s...
gomote list 2>&1 | grep solaris; ls /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/
Check gomote expiry and artifacts
rsc-gotip-solaris-amd64-0 gotip-solaris-amd64 gotip-solaris-amd64 expires in 27m51.222932431s diag.sh fd_poll_runtime.go.orig fd_unix.go.orig iss.json nethttp.assoc.test nethttp.dbg2.test nethttp.dbg.test nethttp.fix.test nethttp.test netpoll.go.orig netpoll_solaris.go.orig portbug2.c portbug.c proposed-fix.diff repro.go rst.go rst.solaris server.go.orig solrepro.go solrepro.solaris stress.sh
Root-caused it. It's a Solaris kernel bug, not a Go bug — and I have a C reproducer and a validated mitigation.
Reproducing it
gomote create gotip-solaris-amd64, push a cross-compiled net/http.test, then run 8 concurrent copies of -test.run '^TestTransportGzip$' -test.count=400. 2–4 of 8 wedge every time, with exactly the CI signature: httptest.Server blocked in Close, one handler goroutine parked in internal/poll.(*FD).Write, no background reader, no client goroutines.
What's actually wrong
At the moment of the hang, netstat shows no trace of the connection — it's fully torn down. I instrumented internal/poll.FD.Write to bound the netpoll wait and then retry the write syscall directly:
STUCKWRITE: fd 8 no write event for 15s; remaining 454925 | pd.user = 4 assocs = 392 events = 337 lastEvents = 1
STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe
pd.user = 4 is POLLOUT — the runtime has the fd associated and is waiting. lastEvents = 1 is POLLIN — the last event Solaris ever delivered for that fd was read-only. And write() returns EPIPE instantly. The kernel simply never reports the fd again.
I reduced it to ~100 lines of C with no Go involved (portbug2.c, on the gomote):
== send buffer filled = no error consumed by read() = no
poll(IN|OUT) -> rc=1 revents=0x5
port_get(OUT) -> events=0x4
write() -> n=-1 errno=32 (Broken pipe)
SO_ERROR -> 131 (Connection reset by peer)
== send buffer filled = no error consumed by read() = yes
read() -> n=-1 errno=131 (Connection reset by peer)
poll(IN|OUT) -> rc=1 revents=0x1
port_get(OUT) -> TIMED OUT after 5s <<< no event, ever
write() -> n=-1 errno=32 (Broken pipe)
SO_ERROR -> 0 (-)
Once a read() has consumed the socket's pending ECONNRESET, Solaris 11.4 reports the fd as POLLIN only — never POLLOUT, POLLERR, or POLLHUP — even though write() fails immediately with EPIPE. poll(2) is just as blind as event ports, so this isn't specific to the event-port path or to netpoll_solaris.go. Whether the send buffer is full is irrelevant; consuming the error via read() is the only precondition.
Maps onto the test exactly: the client resets the connection, net/http's connReader.backgroundRead consumes the error and exits, and the handler's io.Copy(rw, &buf) — armed for POLLOUT — is never woken again.
Two things I ruled out along the way: always calling port_associate in netpollupdate (dropping the old == events early return) makes no difference — 3/8 still wedged; and the isolated Go model with a blocked write plus a concurrent reader passes on Solaris in all FIN/RST variants, because there both modes are armed when the peer dies and the single delivered event carries POLLIN|POLLOUT.
That also explains the March start date without blaming CL 737720 for a bug: the drain added by 11d5284363 (client reads 256 KiB/50 ms before closing) means the server cycles through many wake/write/re-arm rounds, so the RST can land while only POLLIN is armed. The drain widened the window on a pre-existing kernel bug; it didn't introduce one.
Mitigation
net/http already knows the connection is dead — handleReadErrorLocked cancels the request context. It can also unblock the wedged write:
func (cr *connReader) handleReadErrorLocked(err error) {
if cr.conn == nil {
return
}
// A hard read error means the connection is gone, not merely
// half-closed, so abort any response write in flight. ...
if err != nil && err != io.EOF {
if ne, ok := err.(net.Error); !ok || !ne.Timeout() {
cr.conn.rwc.SetWriteDeadline(aLongTimeAgo)
}
}
cr.conn.cancelCtx()
...
EOF and timeouts are excluded so half-close and ReadTimeout semantics don't change. Results:
| wedged | |
|---|---|
| baseline | 2/8, 3/8, 4/8 |
| with fix | 0/8, 0/8, 0/8 |
Full net/http suite on Solaris: identical 16 failures with and without the patch (all TestServeFile-family, from running the binary outside a GOROOT). On linux/amd64, net/http, net/http/httputil, net/http/httptest all pass.
The patch is applied in your tree (src/net/http/server.go, the only modified file) and saved as scratchpad/proposed-fix.diff. Everything else I touched — netpoll.go, netpoll_solaris.go, fd_unix.go, fd_poll_runtime.go — is reverted.
Caveat on scope
This fixes the net/http server, but any Go program on Solaris with concurrent read and write goroutines on one connection can wedge the same way when the peer resets — the reader eats the error and the writer waits forever. A general runtime fix is awkward: at wedge time only POLLOUT is armed, so waking writers on POLLIN events does nothing, and associating POLLIN on write-arm risks a spin loop when unread data sits in the receive buffer. The real fix belongs in Solaris, and portbug2.c is the artifact to send them (cc @golang/solaris, @rorth; it also belongs on the #60152 umbrella).
The gomote has ~28 min left. Want me to draft the issue comment, or keep the instance for more experiments?
You Wed 26 Aug 2026 12:39:43
Is portbug2.c also local?
Assistant Wed 26 Aug 2026 12:39:46
Yes — it was written locally first, then pushed to the buildlet. Both C repros are in the scratchpad:
Ran tools (1 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; ls -l $SP/*.c $SP/prop...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; ls -l $SP/*.c $SP/proposed-fix.diff $SP/stress.sh
List local C repro files
-rw-rw-r-- 1 claude claude 3211 Aug 26 12:28 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/portbug2.c -rw-rw-r-- 1 claude claude 4082 Aug 26 12:27 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/portbug.c -rw-rw-r-- 1 claude claude 1048 Aug 26 12:36 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/proposed-fix.diff -rw-rw-r-- 1 claude claude 399 Aug 26 12:11 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/stress.sh
portbug2.c is the minimal one — the 2×2 matrix (send buffer filled or not × error consumed by read() or not), always closing the peer with SO_LINGER 0. That's the one to attach to a bug report. portbug.c was the earlier version that also varies FIN vs RST; its FIN cases all behave correctly, so it's mostly of historical interest.
Build on Solaris with gcc -o portbug2 portbug2.c -lsocket -lnsl.
Want me to move them somewhere more permanent than the scratchpad?
You Wed 26 Aug 2026 12:40:11
Which SRU is the gomote running?
Assistant Wed 26 Aug 2026 12:40:16
Ran tools (1 Bash)
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'cat /etc/release; echo ---; uname -v; echo ---; pkg info entire 2...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'cat /etc/release; echo ---; uname -v; echo ---; pkg info entire 2>&1 | head -20' 2>&1 | tail -30
Get Solaris SRU version
180000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote288184879/rsc-gotip-solaris-amd64-0.stdout"...
Oracle Solaris 11.4 X86
Copyright (c) 1983, 2026, Oracle and/or its affiliates.
Assembled 09 July 2026
---
11.4.95.226.0
---
Name: entire
Summary: entire incorporation including Support Repository Update
(Oracle Solaris 11.4.95.226.0).
Description: This package constrains system package versions to the same
build. WARNING: Proper system update and correct package
selection depend on the presence of this incorporation.
Removing this package will result in an unsupported system.
For more information see:
https://support.oracle.com/rs?type=doc&id=2433412.1
Category: Meta Packages/Incorporations
State: Installed
Publisher: solaris
Version: 11.4 (Oracle Solaris 11.4.95.226.0)
Branch: 11.4.95.0.0.226.0
Packaging Date: July 10, 2026 at 2:47:09 PM
Last Install Time: March 19, 2019 at 7:56:51 PM
Last Update Time: July 24, 2026 at 10:08:06 PM
Size: 2.52 kB
FMRI: pkg://solaris/entire@11.4-11.4.95.0.0.226.0:20260710T144709Z
# Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote288184879/rsc-gotip-solaris-amd64-0.stdout".Oracle Solaris 11.4 SRU 95 (11.4.95.226.0, branch 11.4.95.0.0.226.0), packaged 2026-07-10, updated on the machine 2026-07-24. Media assembled 2026-07-09, x86.
Worth noting for the bug report: the earliest failure on this issue is 2026-03-28, so the builder was on an older SRU then and the behavior has survived at least a few SRU updates.
You Wed 26 Aug 2026 12:40:51
Does Solaris have a public bug tracker and/or source code?
Assistant Wed 26 Aug 2026 12:41:08
Ran tools (2 WebSearch, 2 Bash)
WebSearch illumos public bug tracker illumos.org issues source illumos-gate
illumos public bug tracker illumos.org issues source illumos-gate
Web search results for query: "illumos public bug tracker illumos.org issues source illumos-gate"
Links: [{"title":"https://www.illumos.org/issues.csv?assigned_to_id=...","url":"https://www.illumos.org/issues.csv?assigned_to_id=14&per_page=25&set_filter=1&sort=assigned_to,status:desc,done_ratio&status_id=*"},{"title":"Overview - illumos gate - illumos","url":"https://www.illumos.org/projects/illumos-gate"},{"title":"Code wiki","url":"https://wiki.illumos.org/plugins/viewsource/viewpagesrc.action?pageId=1146926"},{"title":"Bug #10167: unable to check out illumos-gate on Windows - illumos gate - illumos","url":"https://www.illumos.org/issues/10167"},{"title":"How To Report Problems - illumos - illumos wiki","url":"https://wiki.illumos.org/display/illumos/How+To+Report+Problems"},{"title":"Overview - illumos","url":"https://illumos.org/docs/contributing/"},{"title":"Bug Reporting - illumos","url":"https://illumos.org/docs/user-guide/bugs/"},{"title":"All unresolved bugs - illumos gate - illumos","url":"https://www.illumos.org/projects/illumos-gate/issues?page=6&per_page=250&query_id=4&sort=status,updated_on:desc,tracker"},{"title":"illumos","url":"https://illumos.org/"}]
Based on the search results, here's what I found about the illumos public bug tracker and illumos-gate:
illumos gate is an open-source Unix operating system, and you can report a new issue using the issue tracker at illumos.org.
Important links include a Bug Tracker, along with illumos Source on Gerrit (primary repository) and illumos Source on GitHub (read-only mirror).
The illumos-gate project is accessible at https://www.illumos.org/projects/illumos-gate, where you can view:
- Issues managed by Dan McDonald, Electric Monk, Gordon Ross, Hans Rosenfeld, Joshua M. Clulow, Patrick Mooney, and Robert Mustacchi
For development work, there is an OpenGrok source browser and search engine at src.illumos.org.
To file issues, you need to create an account in the bug tracker, and before creating new issues, you should perform a search to determine if the problem is already known.
REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.WebSearch Oracle Solaris 11.4 report a bug without support contract public bug database
Oracle Solaris 11.4 report a bug without support contract public bug database
Web search results for query: "Oracle Solaris 11.4 report a bug without support contract public bug database"
Links: [{"title":"Previously Documented Bugs That Are Fixed in the Oracle Solaris 11.4 Release - Oracle® Solaris 11.4 Release Notes","url":"https://docs.oracle.com/cd/E37838_01/html/E60973/gmfda.html"},{"title":"Previously Documented Bugs That Are Fixed in the Oracle Solaris 11.2 Release - Oracle® Solaris 11.2 Release Notes","url":"https://docs.oracle.com/cd/E36784_01/html/E36797/gmfda.html"},{"title":"17 Submit a Bug Report - Java","url":"https://docs.oracle.com/javase/10/troubleshoot/submit-bug-report.htm"},{"title":"Troubleshooting System Administration Issues in Oracle ® Solaris 11.4","url":"https://docs.oracle.com/cd/E37838_01/pdf/E61000.pdf"},{"title":"Updating Systems and Adding Software in Oracle Solaris 11.4 E60979-03 May 2025","url":"https://docs.oracle.com/en/operating-systems/solaris/oracle-solaris/11.4/update-sys-add-sw/updating-systems-and-adding-software-oracle-solaris-11.4.pdf"},{"title":"Why Oracle Support site hides bug reports? - Oracle Forums","url":"https://forums.oracle.com/ords/apexds/post/why-oracle-support-site-hides-bug-reports-8111"},{"title":"Oracle Solaris 11.4 | Specs, reviews and EoL info","url":"https://invgate.com/itdb/oracle-solaris-11"},{"title":"docs.oracle.com","url":"https://docs.oracle.com/cd/E26502_01/html/E28978/glnpj.html"},{"title":"preface 13","url":"https://docs.oracle.com/cd/E26505_01/html/E29493/preface-13.html"},{"title":"preface 13","url":"https://docs.oracle.com/cd/E26502_01/html/E28978/preface-13.html"}]
Based on the search results, here's what I found regarding reporting bugs for Oracle Solaris 11.4:
## Bug Reporting and Access
Oracle's Sun System defects have migrated to Oracle's bug database, and customers with support contracts can use My Oracle Support (MOS) to search bug information using BugDB bug IDs.
However, the search results indicate that bug access through Oracle's official channels is **limited to customers with support contracts**. To view information about a bug or enhancement request associated with a Service Request (SR), you can see the Sun Systems Defects Move to Oracle's Bug Database (Doc ID 1501467.1) knowledge article available on MOS.
## Public Bug Database
The search results don't reveal a publicly accessible bug database for Oracle Solaris 11.4 that's available without a support contract. The documentation consistently references My Oracle Support as the primary channel for accessing bug information, which requires a support contract.
If you're looking to report or access bugs without a support contract, you may want to search for alternative channels or community resources, as the official Oracle documentation appears to restrict detailed bug information access to contract holders.
REMINDER: You MUST include the sources above in your response to the user using markdown hyperlinks.Bash head -8 src/runtime/netpoll_solaris.go | grep -i "go:build" ; grep -rn "illumos" src/go/build/syslist.go | head -3; grep...
head -8 src/runtime/netpoll_solaris.go | grep -i "go:build" ; grep -rn "illumos" src/go/build/syslist.go | head -3; grep -rn '"illumos"' src/cmd/dist/build.go | head -3; grep -rn "solaris" src/internal/platform/supported.go | head -5
Check whether illumos shares the solaris netpoll
ugrep: warning: src/go/build/syslist.go: No such file or directory src/cmd/dist/build.go:93: "illumos", src/cmd/dist/build.go:1088: "illumos": true, src/cmd/dist/build.go:1104: return goos == "solaris" || goos == "illumos"
Bash sed -n '1,12p' src/runtime/netpoll_solaris.go | cat -A | grep -n "build" ; echo "---"; sed -n '1,10p' src/runtime/netpol...
sed -n '1,12p' src/runtime/netpoll_solaris.go | cat -A | grep -n "build" ; echo "---"; sed -n '1,10p' src/runtime/netpoll_solaris.go; echo "=== illumos implies solaris tag:"; grep -rn "illumos" src/go/build/build.go | head -5
Show build constraint of netpoll_solaris.go
---
// Copyright 2014 The Go Authors. All rights reserved.
// Use of this source code is governed by a BSD-style
// license that can be found in the LICENSE file.
package runtime
import (
"internal/goarch"
"internal/runtime/atomic"
"unsafe"
=== illumos implies solaris tag:
src/go/build/build.go:1940:// solaris (if GOOS = illumos)
src/go/build/build.go:1962: if ctxt.GOOS == "illumos" && name == "solaris" {
src/go/build/build.go:1993:// if GOOS=illumos, then files with GOOS=solaris are also matched.Source: no for Oracle Solaris, yes for illumos. Oracle Solaris 11.x is closed source — OpenSolaris was discontinued in 2010. The open continuation is illumos, with illumos-gate on Gerrit, a read-only GitHub mirror, and an OpenGrok browser at src.illumos.org. It descends from the same OpenSolaris code, so the sockfs/poll logic is close kin to what's misbehaving here, though 15 years diverged and Oracle's is the one actually running on the builder.
Bug tracker: no for Oracle, yes for illumos. Oracle Solaris defects live in Oracle's internal BugDB, reachable only through My Oracle Support with a support contract; there's no public database and no public submission path. Release notes list fixed bug IDs, but you can't read or file one without a contract. illumos, by contrast, has an open bug tracker (free account, Redmine-based) and reporting guidelines.
Practically, that means the Oracle-side report has to go through someone with a contract — which is an argument for routing it via @rorth, who maintains the Solaris port, rather than filing it ourselves.
One thing worth checking either way: GOOS=illumos also sets the solaris build tag (src/go/build/build.go:1962), so illumos runs the exact same runtime/netpoll_solaris.go. If illumos has the same poll behavior, this isn't an Oracle-only quirk, it's inherited — and it'd be reportable in the open, with a fix that could be read. I can grep the illumos sockfs poll path (so_poll/sopoll, and where so_error is cleared on read) to see whether it has the same hole, or just build and run portbug2.c if there's an illumos machine reachable. Want me to?
You Wed 26 Aug 2026 12:42:45
I don't know whether gotip-illumos-amd64 works but you can tryi t.
Assistant Wed 26 Aug 2026 12:42:48
Ran tools (4 Bash)
Bash gomote create 2>&1 | grep -i illumos
gomote create 2>&1 | grep -i illumos
Check for illumos builder types
120000
(Bash completed with no output)
Bash gomote create 2>&1 | head -60
gomote create 2>&1 | head -60
List available buildlet types
120000
create usage: gomote create [create-opts] <type>
If there's a valid group specified, new instances are
automatically added to the group. If the group in
$GOMOTE_GROUP doesn't exist, and there's no other group
specified, it will be created and new instances will be
added to that group.
Run 'gomote create -list' to see a list of valid builder
types.
Builder types are structured according to the following
format, where the bracketed parts are optional:
[<subrepo>-]<go branch>-<goos>-<goarch>[_<host>][-<mods>*]
Subrepo names always start with 'x_'. Go branch names are
either 'gotip' or 'go<version>' like 'go1.23'. goos and goarch
are the same as the values you'd use in build tags and all
lower-case. The host suffix is optional and you likely do not
need to specify it, but see the full list for what's available.
It's usually just an indicator of the OS version, like '13' to
indicate macOS 13 for darwin/amd64 builders. Mods are specifiers
like 'race' and 'longtest'.
gomotes are set up with the same code used to set up the
environment on the builder except without a Go toolchain.
Subrepo gomotes set up a copy of the subrepo in the workdir,
a full git checkout sync'd to tip-of-tree.
Flags:
-count int
number of instances to create (default 1)
-list
list builder types and exit
-new-group string
also create a new group and add the new instances to it
-setup
set up the instance by pushing GOROOT and building the Go toolchain
-status
print regular status updates while waiting (default true)
-use-golangbuild
disable the installation of build dependencies installed by golangbuild (default true)Bash gomote create -list 2>&1 | grep -i "illumos\|solaris" | head -20
gomote create -list 2>&1 | grep -i "illumos\|solaris" | head -20
Search builder list for illumos
120000
go1.26-illumos-amd64 go1.26-solaris-amd64 go1.27-illumos-amd64 go1.27-solaris-amd64 gotip-illumos-amd64 gotip-solaris-amd64 x_arch-go1.26-solaris-amd64 x_arch-go1.27-solaris-amd64 x_arch-gotip-solaris-amd64 x_benchmarks-go1.26-solaris-amd64 x_benchmarks-go1.27-solaris-amd64 x_benchmarks-gotip-solaris-amd64 x_crypto-go1.26-solaris-amd64 x_crypto-go1.27-solaris-amd64 x_crypto-gotip-solaris-amd64 x_debug-go1.26-solaris-amd64 x_debug-go1.27-solaris-amd64 x_debug-gotip-solaris-amd64 x_image-go1.26-solaris-amd64 x_image-go1.27-solaris-amd64
Bash gomote create gotip-illumos-amd64 2>&1 | tail -3
gomote create gotip-illumos-amd64 2>&1 | tail -3
Create illumos gomote
900000
# still creating gotip-illumos-amd64 (1) after 6s; 0 requests ahead of you # Error running create: failed to create buildlet (1): rpc error: code = Internal desc = gomote creation failed: note that builder has known issue go.dev/issue/67302
illumos gomotes are broken (known issue #67302). Let me read the illumos source instead.
Ran tools (6 Bash)
Bash cd /tmp; curl -sL "https://raw.githubusercontent.com/illumos/illumos-gate/master/usr/src/uts/common/fs/sockfs/sockcommon...
cd /tmp; curl -sL "https://raw.githubusercontent.com/illumos/illumos-gate/master/usr/src/uts/common/fs/sockfs/sockcommon_sops.c" -o sockcommon_sops.c; ls -l sockcommon_sops.c; grep -n "so_poll" sockcommon_sops.c | head
Fetch illumos sockfs source
120000
-rw-rw-r-- 1 claude claude 48078 Aug 26 12:43 sockcommon_sops.c 874:so_poll(struct sonode *so, short events, int anyyet, short *reventsp, 972: so->so_pollev |= SO_POLLEV_IN; 973: *phpp = &so->so_poll_list; 981: so->so_pollev |= SO_POLLEV_IN; 985: *phpp = &so->so_poll_list; 1945: so_poll, /* sop_poll */ Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; sed -n '868,995p' sockcommon_sops.c
cd /tmp; sed -n '868,995p' sockcommon_sops.c
Read illumos so_poll
SO_UNBLOCK_FALLBACK(so);
return (error);
}
int
so_poll(struct sonode *so, short events, int anyyet, short *reventsp,
struct pollhead **phpp)
{
int state = so->so_state, mask;
*reventsp = 0;
/*
* In sockets the errors are represented as input/output events
*/
if (so->so_error != 0 &&
((POLLIN|POLLRDNORM|POLLOUT) & events) != 0) {
*reventsp = (POLLIN|POLLRDNORM|POLLOUT) & events;
return (0);
}
/*
* If the socket is in a state where it can send data
* turn on POLLWRBAND and POLLOUT events.
*/
if ((so->so_mode & SM_CONNREQUIRED) == 0 || (state & SS_ISCONNECTED)) {
/*
* out of band data is allowed even if the connection
* is flow controlled
*/
*reventsp |= POLLWRBAND & events;
if (!SO_SND_FLOWCTRLD(so)) {
/*
* As long as there is buffer to send data
* turn on POLLOUT events
*/
*reventsp |= POLLOUT & events;
}
}
/*
* Turn on POLLIN whenever there is data on the receive queue,
* or the socket is in a state where no more data will be received.
* Also, if the socket is accepting connections, flip the bit if
* there is something on the queue.
*
* We do an initial check for events without holding locks. However,
* if there are no event available, then we redo the check for POLLIN
* events under the lock.
*/
/* Pending connections */
if (!list_is_empty(&so->so_acceptq_list))
*reventsp |= (POLLIN|POLLRDNORM) & events;
/*
* If we're looking for POLLRDHUP, indicate it if we have sent the
* last rx signal for the socket.
*/
if ((events & POLLRDHUP) && (state & SS_SENTLASTREADSIG))
*reventsp |= POLLRDHUP;
/* Data */
/* so_downcalls is null for sctp */
if (so->so_downcalls != NULL && so->so_downcalls->sd_poll != NULL) {
*reventsp |= (*so->so_downcalls->sd_poll)
(so->so_proto_handle, events & SO_PROTO_POLLEV, anyyet,
CRED()) & events;
ASSERT((*reventsp & ~events) == 0);
/* do not recheck events */
events &= ~SO_PROTO_POLLEV;
} else {
if (SO_HAVE_DATA(so))
*reventsp |= (POLLIN|POLLRDNORM) & events;
/* Urgent data */
if ((state & SS_OOBPEND) != 0) {
*reventsp |= (POLLRDBAND | POLLPRI) & events;
}
/*
* If the socket has become disconnected, we set POLLHUP.
* Note that if we are in this state, we will have set POLLIN
* (SO_HAVE_DATA() is true on a disconnected socket), but not
* POLLOUT (SS_ISCONNECTED is false). This is in keeping with
* the semantics of POLLHUP, which is defined to be mutually
* exclusive with respect to POLLOUT but not POLLIN. We are
* therefore setting POLLHUP primarily for the benefit of
* those not polling on POLLIN, as they have no other way of
* knowing that the socket has been disconnected.
*/
mask = SS_SENTLASTREADSIG | SS_SENTLASTWRITESIG;
if ((state & (mask | SS_ISCONNECTED)) == mask)
*reventsp |= POLLHUP;
}
if ((!*reventsp && !anyyet) || (events & POLLET)) {
/* Check for read events again, but this time under lock */
if (events & (POLLIN|POLLRDNORM)) {
mutex_enter(&so->so_lock);
if (SO_HAVE_DATA(so) ||
!list_is_empty(&so->so_acceptq_list)) {
if (events & POLLET) {
so->so_pollev |= SO_POLLEV_IN;
*phpp = &so->so_poll_list;
}
mutex_exit(&so->so_lock);
*reventsp |= (POLLIN|POLLRDNORM) & events;
return (0);
} else {
so->so_pollev |= SO_POLLEV_IN;
mutex_exit(&so->so_lock);
}
}
*phpp = &so->so_poll_list;
}
return (0);
}
/*
* Generic Upcalls
*/
void
so_connected(sock_upper_handle_t sock_handle, sock_connid_t id,
cred_t *peer_cred, pid_t peer_cpid)
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; for f in socknotify.c sockcommon_subr.c sockcommon.c; do curl -sL "https://raw.githubusercontent.com/illumos/il...
cd /tmp; for f in socknotify.c sockcommon_subr.c sockcommon.c; do curl -sL "https://raw.githubusercontent.com/illumos/illumos-gate/master/usr/src/uts/common/fs/sockfs/$f" -o "$f"; done; grep -n "SENTLASTWRITESIG\|SENTLASTREADSIG" *.c | head -30; echo "=== defn ==="; curl -sL "https://raw.githubusercontent.com/illumos/illumos-gate/master/usr/src/uts/common/sys/socketvar.h" -o socketvar.h; grep -n "SENTLAST" socketvar.h
Find SS_SENTLASTWRITESIG usage in illumos
180000
sockcommon_sops.c:927: if ((events & POLLRDHUP) && (state & SS_SENTLASTREADSIG))
sockcommon_sops.c:959: mask = SS_SENTLASTREADSIG | SS_SENTLASTWRITESIG;
sockcommon_subr.c:2219: SS_HAVEOOBDATA|SS_HADOOBDATA|SS_SENTLASTREADSIG|SS_SENTLASTWRITESIG)
socknotify.c:374: if (!(so->so_state & SS_SENTLASTREADSIG)) {
socknotify.c:377: so->so_state |= SS_SENTLASTREADSIG;
socknotify.c:395: if (!(so->so_state & SS_SENTLASTWRITESIG)) {
socknotify.c:397: so->so_state |= SS_SENTLASTWRITESIG;
=== defn ===
317:#define SS_SENTLASTREADSIG 0x01000000 /* last rx signal has been sent */
318:#define SS_SENTLASTWRITESIG 0x02000000 /* last tx signal has been sent */
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; sed -n '355,410p' socknotify.c
cd /tmp; sed -n '355,410p' socknotify.c
Read illumos socket disconnect notification
if (so->so_state & SS_CANTRCVMORE)
(void) i_so_notify_last_rx(so, &pollev, &sigev);
if (sigev != 0)
socket_sendsig(so, sigev);
mutex_exit(&so->so_lock);
if (pollev != 0)
pollwakeup(&so->so_poll_list, pollev);
ASSERT(MUTEX_NOT_HELD(&so->so_lock));
}
/*
* No more data will be coming in, and this will be the last notification
* made.
*/
static int
i_so_notify_last_rx(struct sonode *so, int *pollev, int *sigev)
{
if (!(so->so_state & SS_SENTLASTREADSIG)) {
SOCKET_TIMER_CANCEL(so);
SO_WAKEUP_READER(so);
so->so_state |= SS_SENTLASTREADSIG;
so->so_pollev &= ~SO_POLLEV_IN;
*pollev |= POLLIN|POLLRDNORM|POLLRDHUP;
*sigev |= SOCKETSIG_READ;
return (1);
} else {
return (0);
}
}
/*
* The socket is un-writeable. Make one last notification.
*/
static int
i_so_notify_last_tx(struct sonode *so, int *pollev, int *sigev)
{
if (!(so->so_state & SS_SENTLASTWRITESIG)) {
SO_WAKEUP_WRITER(so);
so->so_state |= SS_SENTLASTWRITESIG;
*pollev |= POLLOUT;
*sigev |= SOCKETSIG_WRITE;
return (1);
} else {
return (0);
}
}
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; grep -n "i_so_notify_last_tx\|i_so_notify_last_rx\|so_notify_disconnected\|so_notify_error" socknotify.c | head...
cd /tmp; grep -n "i_so_notify_last_tx\|i_so_notify_last_rx\|so_notify_disconnected\|so_notify_error" socknotify.c | head -20; echo "=== disconnected/error notifiers ==="; sed -n '/^so_notify_disconnected/,/^}/p' socknotify.c; echo "---- error ----"; sed -n '/^so_notify_error/,/^}/p' socknotify.c
Inspect illumos disconnect/error notifiers
61:static int i_so_notify_last_rx(struct sonode *, int *, int *);
62:static int i_so_notify_last_tx(struct sonode *, int *, int *);
101: (void) i_so_notify_last_tx(so, &pollev, &sigev);
124:so_notify_disconnected(struct sonode *so, boolean_t connfailed, int error)
131: (void) i_so_notify_last_tx(so, &pollev, &sigev);
132: (void) i_so_notify_last_rx(so, &pollev, &sigev);
221:so_notify_error(struct sonode *so)
297: (void) i_so_notify_last_rx(so, &pollev, &sigev);
354: (void) i_so_notify_last_tx(so, &pollev, &sigev);
356: (void) i_so_notify_last_rx(so, &pollev, &sigev);
372:i_so_notify_last_rx(struct sonode *so, int *pollev, int *sigev)
393:i_so_notify_last_tx(struct sonode *so, int *pollev, int *sigev)
=== disconnected/error notifiers ===
so_notify_disconnected(struct sonode *so, boolean_t connfailed, int error)
{
int pollev = 0;
int sigev = 0;
ASSERT(MUTEX_HELD(&so->so_lock));
(void) i_so_notify_last_tx(so, &pollev, &sigev);
(void) i_so_notify_last_rx(so, &pollev, &sigev);
if (IS_KERNEL_SOCKET(so)) {
if (connfailed) {
KSOCKET_CALLBACK(so, disconnected, error);
} else {
KSOCKET_CALLBACK(so, connectfailed, error);
}
mutex_exit(&so->so_lock);
} else {
if (sigev != 0)
socket_sendsig(so, sigev);
mutex_exit(&so->so_lock);
/*
* If we're here because the socket has become disconnected,
* we explicitly set POLLHUP. At the same time, we also clear
* POLLOUT, as POLLOUT and POLLHUP are defined to be mutually
* exclusive with respect to one another.
*/
if (!connfailed)
pollev = (pollev | POLLHUP) & ~POLLOUT;
if (pollev != 0)
pollwakeup(&so->so_poll_list, pollev);
}
sof_sonode_notify_filters(so, (connfailed) ? SOF_EV_CONNECTFAILED :
SOF_EV_DISCONNECTED, error);
ASSERT(MUTEX_NOT_HELD(&so->so_lock));
}
---- error ----
so_notify_error(struct sonode *so)
{
ASSERT(MUTEX_HELD(&so->so_lock));
SO_WAKEUP_WRITER(so);
SO_WAKEUP_READER(so);
if (IS_KERNEL_SOCKET(so)) {
KSOCKET_CALLBACK(so, error, 0);
mutex_exit(&so->so_lock);
} else {
socket_sendsig(so, SOCKETSIG_WRITE|SOCKETSIG_READ);
so->so_pollev &= ~SO_POLLEV_IN;
mutex_exit(&so->so_lock);
pollwakeup(&so->so_poll_list, POLLOUT|POLLIN|POLLRDNORM);
}
ASSERT(MUTEX_NOT_HELD(&so->so_lock));
}
Shell cwd was reset to /home/claude/golang_go_78438Bash gomote list 2>&1 | grep solaris; cd /tmp; curl -sL "https://api.github.com/repos/golang/go/issues/67302" | python3 -c " ...
gomote list 2>&1 | grep solaris; cd /tmp; curl -sL "https://api.github.com/repos/golang/go/issues/67302" | python3 -c " import json,sys d=json.load(sys.stdin); print(d['number'], d['state'], d['title']) print(d['body'][:300])"
Check gomote time and illumos builder issue
rsc-gotip-solaris-amd64-0 gotip-solaris-amd64 gotip-solaris-amd64 expires in 25m31.681049735s 67302 open x/build: add LUCI illumos-amd64 builder There currently isn't a LUCI builder that tests the illumos/amd64 port (other than the misc-compile builder, which tests only that the port compiles). This is the tracking issue for it. The next steps that a builder owner will need to follow to make progress here are documented https://go.dev/wik Shell cwd was reset to /home/claude/golang_go_78438
gomote create gotip-illumos-amd64 fails outright — the type is listed but creation is rejected with a pointer to #67302, which is still just the tracking issue for adding a LUCI illumos builder. There's no illumos machine to test on.
So I read the source instead, and it's worth it — illumos looks not to have this bug, which sharpens the Oracle report considerably.
so_poll() in sockcommon_sops.c matches the Solaris behavior I measured, right up to the last step:
if (so->so_error != 0 &&
((POLLIN|POLLRDNORM|POLLOUT) & events) != 0) {
*reventsp = (POLLIN|POLLRDNORM|POLLOUT) & events;
return (0);
}
...
if ((so->so_mode & SM_CONNREQUIRED) == 0 || (state & SS_ISCONNECTED)) {
...
*reventsp |= POLLOUT & events;
That first clause is my error consumed = no case (revents 0x5). Once read() clears so_error we fall through, and SS_ISCONNECTED is gone after a reset, so POLLOUT is deliberately withheld — exactly the 0x1 I measured on Solaris. But illumos then compensates:
/*
* If the socket has become disconnected, we set POLLHUP.
* Note that if we are in this state, we will have set POLLIN
* ... but not POLLOUT (SS_ISCONNECTED is false). ... We are
* therefore setting POLLHUP primarily for the benefit of
* those not polling on POLLIN, as they have no other way of
* knowing that the socket has been disconnected.
*/
mask = SS_SENTLASTREADSIG | SS_SENTLASTWRITESIG;
if ((state & (mask | SS_ISCONNECTED)) == mask)
*reventsp |= POLLHUP;
Both guard bits are set on the disconnect path — so_notify_disconnected() in socknotify.c calls i_so_notify_last_tx() and i_so_notify_last_rx(), which set SS_SENTLASTWRITESIG and SS_SENTLASTREADSIG — so a re-poll of a reset socket returns POLLHUP. Go's netpoll treats POLLHUP as waking both read and write (netpoll_solaris.go ORs it into both mode computations), so the blocked writer would be woken and the hang wouldn't happen.
That comment is essentially describing our bug victim: a writer that isn't polling for POLLIN and has no other way to learn the socket is dead.
On Oracle Solaris 11.4 SRU 95, POLLHUP never appears — not before the read consumes the error (revents 0x5) and not after (revents 0x1). POLLERR (0x8) and POLLHUP (0x10) are reported by poll(2) regardless of the requested mask, so their absence from those revents values is conclusive, not an artifact of what I asked for.
Report to Oracle, then: a socket whose write() returns EPIPE immediately is reported by poll(2) as POLLIN only — no POLLOUT, no POLLERR, no POLLHUP — once its pending error has been consumed by read(). The common-ancestor code sets POLLHUP in precisely this state and documents why. That's a concrete, checkable claim rather than "poll seems broken."
Two caveats to keep honest in any write-up: the illumos conclusion is code-reading only, since I couldn't build or run it, and Oracle's socket stack has diverged for 15 years, so their code may not resemble so_poll() at all — the value is in the intent the shared ancestor documents, not in a guarantee about Oracle's implementation.
The Solaris gomote has ~25 min left if you want anything else measured on it.
You Wed 26 Aug 2026 13:01:51
Did the illumos code about POLLHUP get added after the fork? And would it help us to ask us for POLLHUP explicitly anyway?
Assistant Wed 26 Aug 2026 13:02:20
Two good questions — let me get data on both.
Ran tools (2 Bash)
Bash cd /tmp; curl -sL "https://api.github.com/repos/illumos/illumos-gate/commits?path=usr/src/uts/common/fs/sockfs/sockcommo...
cd /tmp; curl -sL "https://api.github.com/repos/illumos/illumos-gate/commits?path=usr/src/uts/common/fs/sockfs/sockcommon_sops.c&per_page=100" -o c1.json; python3 -c "
import json
d=json.load(open('c1.json'))
print(len(d))
for c in d:
print(c['commit']['author']['date'][:10], c['sha'][:12], c['commit']['message'].split(chr(10))[0][:100])
"List illumos commits touching sockcommon_sops.c
180000
32
2024-12-11 fcdb3229a31d 17568 remove -Wno-unknown-pragmas
2024-07-23 0250c53ad267 16624 Want support for FD_CLOFORK and friends
2024-02-20 369bcc48d88e 16339 sockfs: variable may be used uninitialized
2022-07-04 15f90b02bdac 14768 retire nca
2022-02-26 bbf215553c72 14443 resection manual pages per IPD4
2021-02-01 907c2824088e 14202 Need direct callbacks from socket upcalls via ksocket
2014-06-20 940f8ece5acb 14199 sendfile compat checks shouldn't be done in so_sendmblk
2019-08-22 78a2e113edb6 9531 Want netstat -u to show PIDs associated with sockets
2015-02-15 a5eb7107f06a 5640 want epoll support
2014-02-24 68846fd00135 4627 POLLHUP not generated for disconnected sockets
2010-08-10 e82bc0ba9649 6972175 assertion failed: tcp->tcp_fin_sent, file: ../../common/inet/tcp/tcp_input.c, line: 4306
2010-07-16 b1cd7879d8fc 6963859 ipcl_conn_create() triggers panic in snv_143 on a machine with solaris10 branded zones
2010-06-18 dd49f1255079 6939100 convert KSSL into a socket filter
2010-06-18 3e95bd4ab92a PSARC/2009/590 Socket Filter Framework
2010-04-21 c0dd49bdd68c PSARC/2010/043 Reliable Datagram Service v3
2009-11-11 bd670b35a010 PSARC/2009/331 IP Datapath Refactoring
2009-07-09 081c0aa8fd8d 6844855 8.72% snv_115 vol25_09_loopback_32bit performance regression x86
2009-05-29 8591a19ae762 6837188 MSG_WAITALL skips over urgent mark
2009-04-28 bbc000e522de 6813959 I/OAT should be removed from TCP and STREAMS
2009-02-13 4117443759eb 6799655 sockets need better handling of STREAMS ioctls
2009-02-19 af89d820d181 6801736 sockets should handle POLLWRBAND
2009-02-14 d68ef20e3fe8 6801576 Race in accept path that can lead to a panic
2009-02-12 de8c4a14ec9a PSARC 2007/670 db_credp update
2009-01-24 3986c91e1475 6796699 iscsi client broken: failed to receive login response
2009-01-21 34dfe6834cf5 6789298 SO_{SND,RCV}TIMEO should be passed down to the protocol
2009-01-20 e5083e819e9d 6794368 iscsi client broken: failed to receive login response
2009-01-07 22238f73378c 6788242 SO_{SND,RCV}TIMEO {set,get}sockopt fail with 32-bit binaries on a 64-bit kernel
2008-12-29 a5adac4d8d51 6786739 xtcp_rcvbuf/bind_bad tests failed with unexpected results.
2008-12-21 f0267584ae28 6787347 dgram_peek sotest fails with unexpected result
2008-12-18 07d13e573541 6785994 sendfilev() can return unexpected errno for non-TCP sockets
2008-12-17 d36be52ed67a 6785041 assertion failed: so_verify_oobstate(so), file: ../../common/fs/sockfs/sockcommon.c, line: 6
2008-12-12 0f1702c52013 PSARC 2007/587 Volo -- Low Latency Socket Framework
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; curl -sL "https://api.github.com/repos/illumos/illumos-gate/commits/68846fd00135" -o k.json; python3 -c " ...
cd /tmp; curl -sL "https://api.github.com/repos/illumos/illumos-gate/commits/68846fd00135" -o k.json; python3 -c "
import json
d=json.load(open('k.json'))
print(d['commit']['author']['date'], '|', d['commit']['author']['name'])
print(d['commit']['message'][:600])
for f in d['files']:
print('---', f['filename'])
if 'sockcommon_sops' in f['filename'] or 'socknotify' in f['filename']:
print(f['patch'][:3000])
"Inspect illumos POLLHUP commit
180000
2014-02-24T17:05:14Z | Bryan Cantrill
4627 POLLHUP not generated for disconnected sockets
Reviewed by: Dan McDonald <danmcd@omniti.com>
Reviewed by: Garrett D'Amore <garrett@damore.org>
Reviewed by: Albert Lee <trisk@nexenta.com>
Approved by: Garrett D'Amore <garrett@damore.org>
--- usr/src/uts/common/fs/sockfs/sockcommon_sops.c
@@ -23,6 +23,10 @@
* Copyright (c) 1999, 2010, Oracle and/or its affiliates. All rights reserved.
*/
+/*
+ * Copyright (c) 2014, Joyent, Inc. All rights reserved.
+ */
+
#include <sys/types.h>
#include <sys/param.h>
#include <sys/systm.h>
@@ -903,7 +907,7 @@ int
so_poll(struct sonode *so, short events, int anyyet, short *reventsp,
struct pollhead **phpp)
{
- int state = so->so_state;
+ int state = so->so_state, mask;
*reventsp = 0;
/*
@@ -966,6 +970,22 @@ so_poll(struct sonode *so, short events, int anyyet, short *reventsp,
if ((state & SS_OOBPEND) != 0) {
*reventsp |= (POLLRDBAND | POLLPRI) & events;
}
+
+ /*
+ * If the socket has become disconnected, we set POLLHUP.
+ * Note that if we are in this state, we will have set POLLIN
+ * (SO_HAVE_DATA() is true on a disconnected socket), but not
+ * POLLOUT (SS_ISCONNECTED is false). This is in keeping with
+ * the semantics of POLLHUP, which is defined to be mutually
+ * exclusive with respect to POLLOUT but not POLLIN. We are
+ * therefore setting POLLHUP primarily for the benefit of
+ * those not polling on POLLIN, as they have no other way of
+ * knowing that the socket has been disconnected.
+ */
+ mask = SS_SENTLASTREADSIG | SS_SENTLASTWRITESIG;
+
+ if ((state & (mask | SS_ISCONNECTED)) == mask)
+ *reventsp |= POLLHUP;
}
if (!*reventsp && !anyyet) {
@@ -1022,7 +1042,13 @@ so_disconnected(sock_upper_handle_t sock_handle, sock_connid_t id, int error)
boolean_t connect_failed;
mutex_enter(&so->so_lock);
- connect_failed = so->so_state & SS_ISCONNECTED;
+
+ /*
+ * If we aren't currently connected, then this isn't a disconnect but
+ * rather a failure to connect.
+ */
+ connect_failed = !(so->so_state & SS_ISCONNECTED);
+
so->so_proto_connid = id;
soisdisconnected(so, error);
so_notify_disconnected(so, connect_failed, error);
--- usr/src/uts/common/fs/sockfs/socknotify.c
@@ -23,6 +23,10 @@
* Copyright (c) 2008, 2010, Oracle and/or its affiliates. All rights reserved.
*/
+/*
+ * Copyright (c) 2014, Joyent, Inc. All rights reserved.
+ */
+
#include <sys/types.h>
#include <sys/param.h>
#include <sys/systm.h>
@@ -103,6 +107,7 @@ so_notify_disconnecting(struct sonode *so)
if (sigev != 0)
socket_sendsig(so, sigev);
mutex_exit(&so->so_lock);
+
if (pollev != 0)
pollwakeup(&so->so_poll_list, pollev);
}
@@ -137,6 +142,16 @@ so_notify_disconnected(struct sonode *so, boolean_t connfailed, int error)
if (sigev != 0)
socket_sendsig(so, sigev);
mutex_exit(&so->so_lock);
+
+ /*
+ * If we're here because the socket has become disconnected,
+ * we explicitly set POLLHUP. At the same time, we also clear
+ * POLLOUT, as POLLOUT and POLLHUP are defined to be mutually
+ * exclusive with respect to one another.
+ */
+ if (!connfailed)
+ pollev = (pollev | POLLHUP) & ~POLLOUT;
+
if (pollev != 0)
pollwakeup(&so->so_poll_list, pollev);
}
Shell cwd was reset to /home/claude/golang_go_78438Now the empirical half of your second question:
Ran tools (3 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/portbug3.c <...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cat > $SP/portbug3.c <<'EOF'
/*
* In the wedged state (peer reset, pending error already consumed by read),
* which event masks can wake a writer on solaris? go.dev/issue/78438
* Build: gcc -o portbug3 portbug3.c -lsocket -lnsl
*/
#include <errno.h>
#include <fcntl.h>
#include <netinet/in.h>
#include <poll.h>
#include <port.h>
#include <signal.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/socket.h>
#include <unistd.h>
static void die(const char *m) { perror(m); exit(2); }
static void decode(short r)
{
printf("0x%04x [", r);
if (r & POLLIN) printf(" POLLIN");
if (r & POLLPRI) printf(" POLLPRI");
if (r & POLLOUT) printf(" POLLOUT");
if (r & POLLERR) printf(" POLLERR");
if (r & POLLHUP) printf(" POLLHUP");
if (r & POLLNVAL) printf(" POLLNVAL");
if (r & POLLRDNORM) printf(" POLLRDNORM");
if (r & POLLRDHUP) printf(" POLLRDHUP");
printf(" ]");
}
/* Returns a socket in the wedged state: peer reset, error consumed by read. */
static int wedged(void)
{
int lfd, cfd, sfd;
struct sockaddr_in sa;
socklen_t salen;
char b[1];
struct linger lg;
lfd = socket(AF_INET, SOCK_STREAM, 0);
memset(&sa, 0, sizeof sa);
sa.sin_family = AF_INET;
sa.sin_addr.s_addr = htonl(INADDR_LOOPBACK);
if (bind(lfd, (struct sockaddr *)&sa, sizeof sa) < 0) die("bind");
if (listen(lfd, 1) < 0) die("listen");
salen = sizeof sa;
getsockname(lfd, (struct sockaddr *)&sa, &salen);
cfd = socket(AF_INET, SOCK_STREAM, 0);
if (connect(cfd, (struct sockaddr *)&sa, salen) < 0) die("connect");
sfd = accept(lfd, NULL, NULL);
close(lfd);
fcntl(sfd, F_SETFL, O_NONBLOCK);
lg.l_onoff = 1; lg.l_linger = 0;
setsockopt(cfd, SOL_SOCKET, SO_LINGER, &lg, sizeof lg);
close(cfd);
sleep(1);
(void) read(sfd, b, 1); /* consumes the pending ECONNRESET */
return sfd;
}
static void try_mask(const char *name, short events)
{
int sfd = wedged(), port, rc;
struct pollfd pfd;
port_event_t pe;
timespec_t ts;
pfd.fd = sfd; pfd.events = events; pfd.revents = 0;
rc = poll(&pfd, 1, 0);
printf(" %-22s poll rc=%d revents=", name, rc);
decode(pfd.revents);
port = port_create();
port_associate(port, PORT_SOURCE_FD, sfd, 0, NULL);
if (port_associate(port, PORT_SOURCE_FD, sfd, events, NULL) < 0)
die("port_associate");
ts.tv_sec = 3; ts.tv_nsec = 0;
memset(&pe, 0, sizeof pe);
rc = port_get(port, &pe, &ts);
if (rc < 0 && errno == ETIME)
printf(" port_get: TIMED OUT\n");
else if (rc < 0)
printf(" port_get: errno=%d\n", errno);
else {
printf(" port_get: ");
decode(pe.portev_events);
printf("\n");
}
close(port);
close(sfd);
}
int main(void)
{
signal(SIGPIPE, SIG_IGN);
printf("state: peer reset, pending error already consumed by read()\n");
try_mask("POLLOUT", POLLOUT);
try_mask("POLLOUT|POLLHUP", POLLOUT | POLLHUP);
try_mask("POLLOUT|POLLERR", POLLOUT | POLLERR);
try_mask("POLLOUT|POLLRDHUP", POLLOUT | POLLRDHUP);
try_mask("POLLOUT|POLLIN", POLLOUT | POLLIN);
return 0;
}
EOF
G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/portbug3.c portbug3.c 2>&1|tail -1; gomote run $G /bin/sh -c 'gcc -o portbug3 portbug3.c -lsocket -lnsl && ./portbug3' 2>&1 | tail -20Test which event masks wake a writer
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote851945261/rsc-gotip-solaris-amd64-0.stdout"...
portbug3.c: In function ‘decode’:
portbug3.c:30:17: error: ‘POLLRDHUP’ undeclared (first use in this function); did you mean ‘POLLHUP’?
30 | if (r & POLLRDHUP) printf(" POLLRDHUP");
| ^~~~~~~~~
| POLLHUP
portbug3.c:30:17: note: each undeclared identifier is reported only once for each function it appears in
portbug3.c: In function ‘main’:
portbug3.c:103:49: error: ‘POLLRDHUP’ undeclared (first use in this function); did you mean ‘POLLHUP’?
103 | try_mask("POLLOUT|POLLRDHUP", POLLOUT | POLLRDHUP);
| ^~~~~~~~~
| POLLHUP
# Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote851945261/rsc-gotip-solaris-amd64-0.stdout".
# Error running run: unable to execute /bin/sh: rpc error: code = Unknown desc = command execution failed: exit status 1Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'grep -n "POLLRDHUP\|POLLHUP\|POLLERR" /usr/include/sys/poll.h' 2>...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'grep -n "POLLRDHUP\|POLLHUP\|POLLERR" /usr/include/sys/poll.h' 2>&1 | tail -15
Check poll.h for POLLRDHUP
180000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote4088295192/rsc-gotip-solaris-amd64-0.stdout"... # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote4088295192/rsc-gotip-solaris-amd64-0.stdout". # Error running run: unable to execute /bin/sh: rpc error: code = Unknown desc = command execution failed: exit status 1
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'ls /usr/include/sys/poll.h /usr/include/poll.h; grep -n "POLL" /u...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'ls /usr/include/sys/poll.h /usr/include/poll.h; grep -n "POLL" /usr/include/sys/poll.h | head -25; echo "rc=$?"' 2>&1 | tail -25
Locate poll constants
180000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3174790346/rsc-gotip-solaris-amd64-0.stdout"... /usr/include/poll.h /usr/include/sys/poll.h 8:#ifndef _SYS_POLL_H 9:#define _SYS_POLL_H 30:#define POLLIN 0x0001 /* fd is readable */ 31:#define POLLPRI 0x0002 /* high priority info at fd */ 32:#define POLLOUT 0x0004 /* fd is writeable (won't block) */ 33:#define POLLRDNORM 0x0040 /* normal data is readable */ 34:#define POLLWRNORM POLLOUT 35:#define POLLRDBAND 0x0080 /* out-of-band data is readable */ 36:#define POLLWRBAND 0x0100 /* out-of-band data is writeable */ 38:#define POLLNORM POLLRDNORM 44:#define POLLERR 0x0008 /* fd has error condition */ 45:#define POLLHUP 0x0010 /* fd has been hung up on */ 46:#define POLLNVAL 0x0020 /* invalid pollfd entry */ 48:#define POLLREMOVE 0x0800 /* remove a cached poll fd from /dev/poll */ 54: * Must be bit-wise distinct from the above POLL flags. 56:#define POLLRDDATA 0x0200 /* Wait for M_DATA; ignore M_PROTO only msgs */ 57:#define POLLNOERR 0x0400 /* Ignore POLLERR conditions */ 59:#define POLLCLOSED 0x8000 /* a (cached) poll fd has been closed */ 123:#endif /* _SYS_POLL_H */ rc=0 # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3174790346/rsc-gotip-solaris-amd64-0.stdout".
Solaris 11.4 has no POLLRDHUP at all. Dropping it:
Ran tools (1 Bash)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; python3 - <<'PY' ...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; python3 - <<'PY'
p="/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/portbug3.c"
s=open(p).read()
s=s.replace('\tif (r & POLLRDHUP) printf(" POLLRDHUP");\n','')
s=s.replace('\ttry_mask("POLLOUT|POLLRDHUP", POLLOUT | POLLRDHUP);\n','')
open(p,'w').write(s)
PY
G=rsc-gotip-solaris-amd64-0; gomote put $G $SP/portbug3.c portbug3.c 2>&1|tail -1; gomote run $G /bin/sh -c 'gcc -o portbug3 portbug3.c -lsocket -lnsl && ./portbug3' 2>&1 | tail -20Run mask experiment
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2738514766/rsc-gotip-solaris-amd64-0.stdout"... state: peer reset, pending error already consumed by read() POLLOUT poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT POLLOUT|POLLHUP poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT POLLOUT|POLLERR poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT POLLOUT|POLLIN poll rc=1 revents=0x0001 [ POLLIN ] port_get: 0x0001 [ POLLIN ] # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2738514766/rsc-gotip-solaris-amd64-0.stdout".
Both questions answered, and the first one is a clean hit.
1. Yes — the POLLHUP code is post-fork illumos, not inherited
68846fd00135, 2014-02-24, Bryan Cantrill: "4627 POLLHUP not generated for disconnected sockets". It added exactly the block I quoted to so_poll(), plus the pollev = (pollev | POLLHUP) & ~POLLOUT line in so_notify_disconnected(), and stamped both files with a fresh Joyent copyright. illumos forked in 2010, so Oracle Solaris never had this — it's not a regression on their side, it's a fix they never got.
That reframes the report substantially. It's not "your poll is broken and your own ancestor code proves it." It's "illumos hit this same defect and fixed it in 2014; here's the fix, and here's a 100-line C program showing Solaris 11.4 SRU 95 still has it." Also worth noting from the same commit: it fixed an inverted connect_failed test in so_disconnected() (so->so_state & SS_ISCONNECTED → !(so->so_state & SS_ISCONNECTED)), so the disconnect path was misclassifying disconnects as connect failures. Oracle's code may or may not carry that same inversion — worth flagging as something to look at while they're in there.
2. No — asking for POLLHUP explicitly does not help
Measured directly in the wedged state:
state: peer reset, pending error already consumed by read()
POLLOUT poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT
POLLOUT|POLLHUP poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT
POLLOUT|POLLERR poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT
POLLOUT|POLLIN poll rc=1 revents=0x0001 [ POLLIN ] port_get: 0x0001 [ POLLIN ]
As expected from POSIX: POLLHUP and POLLERR are always reported in revents when set, so requesting them changes nothing — they aren't being masked out, they're never generated. Same through port_associate.
POLLRDHUP isn't an option either: Solaris 11.4's /usr/include/sys/poll.h doesn't define it at all (illumos only got it in 2015 with epoll support, a5eb7107f06a).
POLLIN is the one thing that fires. So a netpoll workaround does exist mechanically — arm POLLIN|POLLOUT for writers — but I'd argue against it, for the reason I flagged earlier: POLLIN on a socket with unread data the application isn't reading is level-triggered and permanent, so writer arms → immediate event → write returns EAGAIN → re-arm → immediate event, a genuine spin. An HTTP server streaming a large response to a client that pipelines requests hits exactly that. The net/http SetWriteDeadline mitigation (0/8 wedged across three runs) stays the better Go-side answer, with the real fix being Oracle porting illumos 4627.
The gomote has ~15 min left. portbug3.c is in the scratchpad alongside the others.
You Wed 26 Aug 2026 13:05:32
Write out a report.md to use as an issue comment that the maintainers can send back to Oracle. Make it concise but complete. Also write and commit the net/http fix to stop the flakes.
Assistant Wed 26 Aug 2026 13:06:49
Ran tools (2 Bash)
Bash cd /home/claude/golang_go_78438/go && git branch --show-current && git status --porcelain && grep -n "oneConnListener" s...
cd /home/claude/golang_go_78438/go && git branch --show-current && git status --porcelain && grep -n "oneConnListener" src/net/http/serve_test.go | head -3 && grep -n '^import\|"errors"' src/net/http/server.go | head -5
Check repo state and helpers
master
M src/net/http/server.go
52:type oneConnListener struct {
56:func (l *oneConnListener) Accept() (c net.Conn, err error) {
67:func (l *oneConnListener) Close() error {
9:import (
14: "errors"Bash git checkout src/net/http/server.go && python3 - <<'PY' ...
git checkout src/net/http/server.go && python3 - <<'PY'
p='src/net/http/server.go'
s=open(p).read()
old='''// handleReadErrorLocked is called whenever a Read from the client returns a
// non-nil error.
//
// The provided non-nil err is almost always io.EOF or a "use of
// closed network connection". In any case, the error is not
// particularly interesting, except perhaps for debugging during
// development. Any error means the connection is dead and we should
// down its context.
//
// The caller must hold connReader.mu.
func (cr *connReader) handleReadErrorLocked(_ error) {
if cr.conn == nil {
return
}
cr.conn.cancelCtx()'''
new='''// handleReadErrorLocked is called whenever a Read from the client returns a
// non-nil error.
//
// The provided non-nil err is almost always io.EOF or a "use of
// closed network connection". Any error means the connection is dead and we
// should down its context. An error that is neither a clean EOF nor an
// expired read deadline also means the connection is dead in both
// directions, so any response write still in flight is aborted.
//
// The caller must hold connReader.mu.
func (cr *connReader) handleReadErrorLocked(err error) {
if cr.conn == nil {
return
}
// A read error that is not a clean EOF (a half close, after which the
// handler may still legitimately write) and not a read deadline means
// the connection is gone, so there is no point in the handler
// continuing to write a response. Unblock any write in flight.
//
// This matters because on some systems the poller never reports the
// socket as writable again once a read has consumed its pending error,
// so a handler blocked writing a large response would otherwise block
// forever. See go.dev/issue/78438.
if !errors.Is(err, io.EOF) {
if ne, ok := err.(net.Error); !ok || !ne.Timeout() {
cr.conn.rwc.SetWriteDeadline(aLongTimeAgo)
}
}
cr.conn.cancelCtx()'''
assert old in s
open(p,'w').write(s.replace(old,new,1))
PY
git diff --statApply final fix
Updated 1 path from the index src/net/http/server.go | 24 +++++++++++++++++++----- 1 file changed, 19 insertions(+), 5 deletions(-)
Now a regression test that works on every platform, using net.Pipe (whose deadlines are real) plus an injectable read error:
Ran tools (3 Bash)
Bash sed -n '45,75p' src/net/http/serve_test.go
sed -n '45,75p' src/net/http/serve_test.go
Inspect oneConnListener
"syscall"
"testing"
"testing/synctest"
"time"
)
type dummyAddr string
type oneConnListener struct {
conn net.Conn
}
func (l *oneConnListener) Accept() (c net.Conn, err error) {
c = l.conn
if c == nil {
err = io.EOF
return
}
err = nil
l.conn = nil
return
}
func (l *oneConnListener) Close() error {
return nil
}
func (l *oneConnListener) Addr() net.Addr {
return dummyAddr("test-address")
}
func (a dummyAddr) Network() string {Bash cat >> src/net/http/serve_test.go <<'EOF' ...
cat >> src/net/http/serve_test.go <<'EOF'
// injectErrConn serves a canned request and then blocks in Read until a read
// error is injected. Writes and deadlines are handled by the embedded
// net.Pipe end, whose deadlines interrupt a blocked Write.
type injectErrConn struct {
net.Conn // a net.Pipe end
req []byte // request bytes to serve before blocking
readErr chan error // receives the error Read should return
}
func (c *injectErrConn) Read(p []byte) (int, error) {
if len(c.req) > 0 {
n := copy(p, c.req)
c.req = c.req[n:]
return n, nil
}
return 0, <-c.readErr
}
// TestServerAbortsWriteOnConnReadError checks that a hard read error on a
// connection unblocks a handler that is blocked writing a response. Without
// it, a handler streaming a response to a client that has gone away can block
// forever on systems where the poller stops reporting the socket as writable
// once a read has consumed its pending error. See go.dev/issue/78438.
func TestServerAbortsWriteOnConnReadError(t *testing.T) {
setParallel(t)
defer afterTest(t)
serverConn, clientConn := net.Pipe()
defer clientConn.Close()
conn := &injectErrConn{
Conn: serverConn,
req: []byte("GET / HTTP/1.1\r\nHost: example.com\r\n\r\n"),
readErr: make(chan error, 1),
}
writing := make(chan struct{})
writeErr := make(chan error, 1)
handler := HandlerFunc(func(w ResponseWriter, r *Request) {
// Nothing ever reads clientConn, so this blocks. net.Pipe is
// unbuffered, so one write is enough.
close(writing)
_, err := w.Write(make([]byte, 1<<20))
writeErr <- err
})
ls := &oneConnListener{conn}
go Serve(ls, handler)
<-writing
// The connection is gone in both directions. The handler must not stay
// blocked in Write.
conn.readErr <- &net.OpError{Op: "read", Err: syscall.ECONNRESET}
select {
case err := <-writeErr:
if err == nil {
t.Fatal("handler write succeeded; want error")
}
case <-time.After(5 * time.Second):
t.Fatal("handler still blocked in Write 5s after the connection died")
}
}
EOF
../bin/go test -count=1 -run TestServerAbortsWriteOnConnReadError -v net/http 2>&1 | tail -15Add and run the regression test
300000
/bin/bash: line 133: ../bin/go: No such file or directory
Bash cd /home/claude/golang_go_78438/go/src && ../bin/go test -count=1 -run TestServerAbortsWriteOnConnReadError -v net/http ...
cd /home/claude/golang_go_78438/go/src && ../bin/go test -count=1 -run TestServerAbortsWriteOnConnReadError -v net/http 2>&1 | tail -15
Run the new test
300000
=== RUN TestServerAbortsWriteOnConnReadError --- PASS: TestServerAbortsWriteOnConnReadError (0.00s) PASS ok net/http 0.004s
Now the crucial check — does the test actually catch the bug?
Ran tools (1 Bash)
Bash cd /home/claude/golang_go_78438/go; cp src/net/http/server.go /tmp/server.go.fixed; git checkout src/net/http/server.go;...
cd /home/claude/golang_go_78438/go; cp src/net/http/server.go /tmp/server.go.fixed; git checkout src/net/http/server.go; cd src && ../bin/go test -count=1 -run TestServerAbortsWriteOnConnReadError -v net/http 2>&1 | tail -8; cd .. ; cp /tmp/server.go.fixed src/net/http/server.go; echo "=== fix restored ==="; git diff --stat
Verify test fails without the fix
300000
Updated 1 path from the index
=== RUN TestServerAbortsWriteOnConnReadError
serve_test.go:8209: handler still blocked in Write 5s after the connection died
--- FAIL: TestServerAbortsWriteOnConnReadError (5.03s)
FAIL
FAIL net/http 5.028s
FAIL
=== fix restored ===
src/net/http/serve_test.go | 64 ++++++++++++++++++++++++++++++++++++++++++++++
src/net/http/server.go | 24 +++++++++++++----
2 files changed, 83 insertions(+), 5 deletions(-)The test fails without the fix and passes with it, on Linux. Now the full verification sweep:
Ran tools (3 Bash)
Bash cd /home/claude/golang_go_78438/go/src && ../bin/go vet net/http 2>&1 | tail -5; ../bin/gofmt -l net/http/server.go net/...
cd /home/claude/golang_go_78438/go/src && ../bin/go vet net/http 2>&1 | tail -5; ../bin/gofmt -l net/http/server.go net/http/serve_test.go; echo "=== tests ==="; ../bin/go test -count=1 net/http net/http/httptest net/http/httputil net/http/cgi net/http/fcgi 2>&1 | tail -10
Vet, gofmt, and run tests on linux
900000
=== tests === ok net/http 11.477s ok net/http/httptest 0.029s ok net/http/httputil 0.627s ok net/http/cgi 0.052s ok net/http/fcgi 0.204s
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; GOOS=solaris GOARCH=am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; GOOS=solaris GOARCH=amd64 ../bin/go test -c -o $SP/nethttp.final.test net/http 2>&1|tail -3; gomote list 2>&1 | grep solaris
Cross-build final binary and check gomote
300000
rsc-gotip-solaris-amd64-0 gotip-solaris-amd64 gotip-solaris-amd64 expires in 25m15.5443718s
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0 gomote put $G $SP/nethttp.final.test nethttp.final.test 2>&1|tail -1 gomote run $G /bin/sh -c 'chmod +x nethttp.final.test; ./nethttp.final.test -test.run "^TestServerAbortsWriteOnConnReadError$" -test.v 2>&1 | tail -4; ./stress.sh nethttp.final.test final' 2>&1 | tail -8
Verify final fix on solaris
900000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3982427909/rsc-gotip-solaris-amd64-0.stdout"... === RUN TestServerAbortsWriteOnConnReadError --- PASS: TestServerAbortsWriteOnConnReadError (0.00s) PASS RESULT final: 0 of 8 wedged # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3982427909/rsc-gotip-solaris-amd64-0.stdout".
Verified on Solaris. Committing:
Ran tools (2 Bash)
Bash cd /home/claude/golang_go_78438/go && git checkout -b fix-78438-write-hang 2>&1 | tail -1 && git add src/net/http/server...
cd /home/claude/golang_go_78438/go && git checkout -b fix-78438-write-hang 2>&1 | tail -1 && git add src/net/http/server.go src/net/http/serve_test.go && git commit -q -F - <<'EOF' && git log --oneline -1 net/http: abort in-flight response writes when the connection dies connReader.handleReadErrorLocked already cancels the connection's context when a read from the client fails, but a handler blocked writing a response is left to discover the dead connection on its own, via the write failing. On Oracle Solaris that never happens: once a read has consumed a socket's pending error, poll(2) reports the fd as POLLIN only, never POLLOUT, POLLERR or POLLHUP, even though write(2) returns EPIPE immediately. A handler blocked writing a large response therefore blocks forever, and with it the connection and anything waiting on it, such as httptest.Server.Close. That is what wedges TestTransportGzip/h1 on the solaris-amd64 builders: the client closes the connection mid-response, the server's background read consumes the ECONNRESET and exits, and the handler's io.Copy of a 1 MiB body never returns. Unblock the write instead of waiting for the poller. A read error that is neither a clean EOF nor an expired read deadline means the connection is gone in both directions, so a response write still in flight is pointless and is now aborted with a write deadline. A clean EOF is left alone, so a client that half closes can still receive a response. The Solaris behavior is an OS bug; illumos fixed the same defect in 2014 (illumos 4627, "POLLHUP not generated for disconnected sockets"). This change also makes such a handler fail promptly rather than leak on any platform where a peer dies without the write ever being reported ready. Fixes #78438 EOF
Branch and commit the fix
Switched to a new branch 'fix-78438-write-hang' 7973707bea net/http: abort in-flight response writes when the connection dies
Bash git commit -q --amend -F - <<'EOF' && git log -1 --format='%s%n%n%b' ...
git commit -q --amend -F - <<'EOF' && git log -1 --format='%s%n%n%b' net/http: abort in-flight response writes when the connection dies connReader.handleReadErrorLocked already cancels the connection's context when a read from the client fails, but a handler blocked writing a response is left to discover the dead connection on its own, via the write failing. On Oracle Solaris that never happens: once a read has consumed a socket's pending error, poll(2) reports the fd as POLLIN only, never POLLOUT, POLLERR or POLLHUP, even though write(2) returns EPIPE immediately. A handler blocked writing a large response therefore blocks forever, and with it the connection and anything waiting on it, such as httptest.Server.Close. That is what wedges TestTransportGzip/h1 on the solaris-amd64 builders: the client closes the connection mid-response, the server's background read consumes the ECONNRESET and exits, and the handler's io.Copy of a 1 MiB body never returns. Unblock the write instead of waiting for the poller. A read error that is neither a clean EOF nor an expired read deadline means the connection is gone in both directions, so a response write still in flight is pointless and is now aborted with a write deadline. A clean EOF is left alone, so a client that half closes can still receive a response. The Solaris behavior is an OS bug; illumos fixed the same defect in 2014 (illumos 4627, "POLLHUP not generated for disconnected sockets"). This change also makes such a handler fail promptly rather than leak on any platform where a peer dies without the write ever being reported ready. Fixes #78438 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Av53TbK6vnxpEPWozrtpr7 EOF
Amend commit with trailers
net/http: abort in-flight response writes when the connection dies connReader.handleReadErrorLocked already cancels the connection's context when a read from the client fails, but a handler blocked writing a response is left to discover the dead connection on its own, via the write failing. On Oracle Solaris that never happens: once a read has consumed a socket's pending error, poll(2) reports the fd as POLLIN only, never POLLOUT, POLLERR or POLLHUP, even though write(2) returns EPIPE immediately. A handler blocked writing a large response therefore blocks forever, and with it the connection and anything waiting on it, such as httptest.Server.Close. That is what wedges TestTransportGzip/h1 on the solaris-amd64 builders: the client closes the connection mid-response, the server's background read consumes the ECONNRESET and exits, and the handler's io.Copy of a 1 MiB body never returns. Unblock the write instead of waiting for the poller. A read error that is neither a clean EOF nor an expired read deadline means the connection is gone in both directions, so a response write still in flight is pointless and is now aborted with a write deadline. A clean EOF is left alone, so a client that half closes can still receive a response. The Solaris behavior is an OS bug; illumos fixed the same defect in 2014 (illumos 4627, "POLLHUP not generated for disconnected sockets"). This change also makes such a handler fail promptly rather than leak on any platform where a peer dies without the write ever being reported ready. Fixes #78438 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Av53TbK6vnxpEPWozrtpr7
Now the report:
Ran tools (1 Write, 1 Bash, 1 SendUserFile)
Write /home/claude/golang_go_78438/report.md
/home/claude/golang_go_78438/report.md
# net/http: TestTransportGzip/h1 failures — root cause
Short version: this is an Oracle Solaris kernel bug in `poll(2)`, not a Go bug.
illumos fixed the same defect in 2014. The section below headed **For Oracle**
is self-contained and can be forwarded as-is.
## What the failures actually are
Every entry watchflakes collected is `[SKIP]` with only `=== RUN/PAUSE/CONT`, so
the issue body shows nothing useful. The real failure in each case is a
package-level timeout of `net/http` with `TestTransportGzip/h1` as the last test
running.
All five solaris-amd64 failures (2026-03-28, 05-19, 07-13, 07-15, 08-21) have an
identical shape:
```
httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0x... 127.0.0.1:47089 in state active
panic: test timed out after 10m0s
running tests:
TestTransportGzip/h1 (9m58s)
```
with exactly two relevant goroutines left: the handler parked in
`internal/poll.(*FD).Write` (the `defer io.Copy(rw, &buf)` at
`transport_test.go:1223`, writing the buffered 1 MiB body), and the test cleanup
in `httptest.Server.Close` waiting on it. `Server.Close` never force-closes
`StateActive` connections, so it waits forever. The test body itself had already
passed.
The `connReader.backgroundRead` goroutine, which `net/http` starts on the same
fd for a body-less request (`server.go:2122`), is absent from every dump — it had
already returned. Its only exit for this connection is a read error. So the read
side of the fd saw the connection die while the write side did not.
The one linux-ppc64le failure (2026-05-18) is unrelated — no blocked handler, the
client wedged in `Transport.getConn` with a dial goroutine that had not run for
17 minutes. That is #78576, and this issue's query is picking it up incidentally.
## Reproducing
On a `gotip-solaris-amd64` gomote, push a cross-compiled `net/http.test` and run
eight concurrent copies:
```sh
for i in 1 2 3 4 5 6 7 8; do
./nethttp.test -test.run '^TestTransportGzip$' -test.count=400 -test.timeout=3m &
done; wait
```
Two to four of the eight wedge on every attempt, with the CI signature above.
## Diagnosis
At the moment of the hang the connection is **completely absent from
`netstat`** — both endpoints are torn down — yet the goroutine is still parked in
netpoll.
Instrumenting `internal/poll.FD.Write` to bound the wait and then retry the
write syscall directly:
```
STUCKWRITE: fd 8 no write event for 15s; remaining 454925 | pd.user = 4 assocs = 392 events = 337 lastEvents = 1
STUCKWRITE: fd 8 retry returned n = -1 err = broken pipe
```
- `pd.user = 4` is `POLLOUT`: the runtime has the fd associated and is waiting.
- `lastEvents = 1` is `POLLIN`: the last event Solaris ever delivered for the fd
was read-only — no `POLLOUT`, `POLLERR` or `POLLHUP`.
- `write(2)` returns `EPIPE` instantly.
So the socket is dead, the write would fail immediately, and the kernel simply
never reports the fd again.
Two things ruled out on the Go side: removing the `old == events` early return in
`netpollupdate` so every arm calls `port_associate` changes nothing (3/8 still
wedged), and an isolated Go model — blocked write plus concurrent reader, peer
closed with FIN or RST — passes on Solaris in all variants, because there both
modes are armed when the peer dies and the single delivered event carries
`POLLIN|POLLOUT`.
## For Oracle
**Oracle Solaris 11.4 SRU 95 (`11.4.95.226.0`), x86.**
Once a TCP socket has been reset by its peer *and* the pending socket error has
been consumed by a `read(2)`, `poll(2)` reports the fd as `POLLIN` only. It
reports neither `POLLOUT` nor `POLLERR` nor `POLLHUP`, even though `write(2)` on
that fd returns `EPIPE` immediately without blocking. A thread waiting for the
socket to become writable is therefore never woken, and cannot learn that the
connection is gone. Event ports (`port_associate`/`port_getn`) behave the same
way, as expected, since they report poll events.
`POLLERR` and `POLLHUP` are reported in `revents` regardless of the requested
event mask, so their absence is not an artifact of what was requested. Whether
the send buffer is full makes no difference. `POLLRDHUP` is not defined on
Solaris 11.4, so it is not an alternative.
Reproducer (`portbug2.c`, attached; `gcc -o portbug2 portbug2.c -lsocket -lnsl`).
It creates a loopback TCP connection, optionally fills the send buffer, closes
the peer with `SO_LINGER 0` to force a reset, optionally consumes the error with
`read(2)`, then polls:
```
== send buffer filled = no error consumed by read() = no
poll(IN|OUT) -> rc=1 revents=0x5
port_get(OUT) -> events=0x4
write() -> n=-1 errno=32 (Broken pipe)
SO_ERROR -> 131 (Connection reset by peer)
== send buffer filled = no error consumed by read() = yes
read() -> n=-1 errno=131 (Connection reset by peer)
poll(IN|OUT) -> rc=1 revents=0x1
port_get(OUT) -> TIMED OUT after 5s <<< no event, ever
write() -> n=-1 errno=32 (Broken pipe)
SO_ERROR -> 0 (-)
```
The second block is the bug: `revents = POLLIN` only, no event from the port,
`write()` fails immediately. Adding `POLLHUP` or `POLLERR` to the requested mask
does not change the result:
```
state: peer reset, pending error already consumed by read()
POLLOUT poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT
POLLOUT|POLLHUP poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT
POLLOUT|POLLERR poll rc=0 revents=0x0000 [ ] port_get: TIMED OUT
POLLOUT|POLLIN poll rc=1 revents=0x0001 [ POLLIN ] port_get: 0x0001 [ POLLIN ]
```
**illumos fixed this exact defect** in commit
[`68846fd00135`](https://github.com/illumos/illumos-gate/commit/68846fd00135)
(2014-02-24), *"4627 POLLHUP not generated for disconnected sockets"*. In
`so_poll()` in `usr/src/uts/common/fs/sockfs/sockcommon_sops.c` the pending-error
clause reports `POLLIN|POLLRDNORM|POLLOUT` while `so_error != 0`, matching the
first block above. Once the error is consumed, `POLLOUT` is deliberately withheld
because `SS_ISCONNECTED` is clear — matching the second block. illumos then
compensates:
```c
/*
* If the socket has become disconnected, we set POLLHUP.
* Note that if we are in this state, we will have set POLLIN
* (SO_HAVE_DATA() is true on a disconnected socket), but not
* POLLOUT (SS_ISCONNECTED is false). This is in keeping with
* the semantics of POLLHUP, which is defined to be mutually
* exclusive with respect to POLLOUT but not POLLIN. We are
* therefore setting POLLHUP primarily for the benefit of
* those not polling on POLLIN, as they have no other way of
* knowing that the socket has been disconnected.
*/
mask = SS_SENTLASTREADSIG | SS_SENTLASTWRITESIG;
if ((state & (mask | SS_ISCONNECTED)) == mask)
*reventsp |= POLLHUP;
```
Both guard bits are set on the disconnect path by `so_notify_disconnected()` in
`socknotify.c`, which the same commit also changed to set `POLLHUP` and clear
`POLLOUT` in its `pollwakeup()`. The commit additionally corrected an inverted
`connect_failed` test in `so_disconnected()` (`so->so_state & SS_ISCONNECTED` →
`!(so->so_state & SS_ISCONNECTED)`), which had been misclassifying disconnects as
connect failures; that inversion may be present in the Solaris sources too and is
worth checking at the same time.
This is post-fork illumos work, so Solaris would not have inherited it — it is a
missing fix rather than a regression.
## Impact
Any Go program on Solaris with one goroutine reading and another writing the same
connection can wedge this way when the peer resets: the reader consumes the
error and exits, and the writer blocks forever holding the connection. A `net/http`
server streaming a large response to a client that goes away is the common case,
and the handler goroutine and connection then leak for the life of the process.
The flake in this issue is that pattern reproduced in a test.
The window widened noticeably after CL 737720 (`11d5284363`, 2026-01-21), which
made `Response.Body.Close` drain up to 256 KiB / 50 ms before closing. The extra
wake/write/re-arm cycles make it much more likely the reset lands while only
`POLLIN` is armed, so the delivered event is `POLLIN`-only. That CL widened an
existing kernel bug; it did not introduce one. It matches the first failure
appearing in March.
## Go-side mitigation
`net/http` already knows the connection is dead — `handleReadErrorLocked` cancels
the connection context on a read error. It can also unblock the wedged write, by
setting a write deadline in the past when the read error is neither a clean EOF
(a half close, after which a handler may still legitimately write) nor an expired
read deadline:
```
wedged out of 8 stress runs
before 2, 3, 4
after 0, 0, 0, 0
```
Full `net/http` suite on Solaris is unchanged by the patch (same 16 pre-existing
failures with and without, all from running the test binary outside a GOROOT), and
`net/http`, `net/http/httptest`, `net/http/httputil`, `net/http/cgi`,
`net/http/fcgi` all pass on linux/amd64. A portable regression test using
`net.Pipe` and an injected read error fails without the change and passes with it.
A runtime-level fix was considered and rejected. In the wedged state only
`POLLOUT` is armed, so waking writers on `POLLIN` events does nothing; the only
mask that fires is `POLLIN`, and arming `POLLIN` alongside `POLLOUT` for writers
would spin, because `POLLIN` on a socket with unread data the application is not
reading is level-triggered and permanent. The real fix belongs in Solaris.
## Notes
- The illumos conclusion is from reading source only. `gomote create
gotip-illumos-amd64` is rejected (#67302 is still the tracking issue for adding
that builder), so this was not verified on a running illumos system. Oracle's
socket stack has also diverged for 15 years and may not resemble `so_poll()`;
the value here is the documented intent of the shared ancestor plus the
measured Solaris behavior, not a claim about Oracle's implementation.
- Oracle Solaris has no public source and no public bug tracker — defects live in
Oracle's internal BugDB, reachable only through My Oracle Support with a
support contract — so this needs to go through someone with a contract.
cc @golang/solaris @rorth
- Related umbrella: #60152.
File created successfully at: /home/claude/golang_go_78438/report.md (file state is current in your context — no need to Read it back)
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cp $SP/portbug2.c $SP/...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; cp $SP/portbug2.c $SP/portbug3.c /home/claude/golang_go_78438/; cd /home/claude/golang_go_78438 && ls; echo "=== commit ==="; cd go && git log --oneline -1 && git status --porcelain
Stage report and repros, confirm clean tree
env.sh go portbug2.c portbug3.c report.md === commit === 3f3c7a57ca net/http: abort in-flight response writes when the connection dies
SendUserFile
Issue comment for #78438 — the "For Oracle" section is self-contained and forwardable. portbug2.c is the attachment it references.
[ "/home/claude/golang_go_78438/report.md", "/home/claude/golang_go_78438/portbug2.c" ]
normal
2 files delivered to user. /home/claude/golang_go_78438/report.md → file_uuid: 3840706f-c11a-4cdb-8a33-188113fdd420 /home/claude/golang_go_78438/portbug2.c → file_uuid: 12530502-8182-4185-8aad-047452baceb8
Both done.
report.md
At /home/claude/golang_go_78438/report.md. Structured so the For Oracle section stands alone and can be forwarded without editing — it names the exact SRU, states the bug in one paragraph, shows the portbug2.c output with the failing case marked, proves POLLHUP/POLLERR aren't merely being masked out, and points at illumos 68846fd00135 with the code and the reasoning behind it. The surrounding sections cover the CI evidence, the reproducer, the instrumentation, what was ruled out, impact, and the Go-side mitigation. portbug2.c and portbug3.c are alongside it for attaching.
I also flagged the inverted connect_failed test that the same illumos commit fixed — worth Oracle checking while they're in that code.
The fix
Committed on branch fix-78438-write-hang (branched off master, which was untouched):
3f3c7a57ca net/http: abort in-flight response writes when the connection dies
src/net/http/serve_test.go | 64 ++++++++++++++++++++++++++++++++++++++++++++++
src/net/http/server.go | 24 +++++++++++++----
handleReadErrorLocked now takes the error it was already being passed and discarding, and sets SetWriteDeadline(aLongTimeAgo) when the error is neither a clean io.EOF nor a net.Error timeout. Half-close keeps working (a client that half-closes can still get a response) and ReadTimeout semantics are unchanged.
The test is the part I'd point a reviewer at first: net.Pipe gives real deadlines, so a fake conn that serves a canned request and then blocks in Read until an error is injected reproduces the shape without needing the OS bug. It fails on Linux without the fix (handler still blocked in Write 5s after the connection died) and passes with it — so it's a genuine regression test everywhere, not just documentation.
Verification:
| result | |
|---|---|
| new test, Linux, without fix | FAIL after 5s |
| new test, Linux + Solaris, with fix | PASS |
| Solaris stress, before | 2/8, 3/8, 4/8 wedged |
| Solaris stress, after | 0/8 |
full net/http suite, Solaris |
identical 16 pre-existing failures either way |
net/http + httptest/httputil/cgi/fcgi, Linux |
all ok |
go vet, gofmt |
clean |
Two things worth your judgment before this becomes a CL. The commit message says Fixes #78438, but the underlying bug is Oracle's and will outlive this change — you may want Updates instead, with the issue staying open against Solaris. And the commit carries the Co-Authored-By/Claude-Session trailers my instructions require, which aren't Gerrit convention; strip them if they'd trip up the CL.
You Wed 26 Aug 2026 13:17:21
What is a link I can add for the illumos change
Assistant Wed 26 Aug 2026 13:17:27
Ran tools (1 Bash, 1 WebFetch)
Bash cd /tmp; python3 -c " ...
cd /tmp; python3 -c "
import json
d=json.load(open('k.json'))
print('full sha:', d['sha'])
print('html_url:', d['html_url'])
"; echo "=== illumos.org issue 4627 ==="; curl -s -o /dev/null -w "%{http_code}\n" https://www.illumos.org/issues/4627Get full SHA and check illumos issue page
120000
full sha: 68846fd00135fb0b10944e7806025cbefcfd6546 html_url: https://github.com/illumos/illumos-gate/commit/68846fd00135fb0b10944e7806025cbefcfd6546 === illumos.org issue 4627 === 200 Shell cwd was reset to /home/claude/golang_go_78438
WebFetch https://www.illumos.org/issues/4627
https://www.illumos.org/issues/4627
Reproduce the issue title, status, author, dates, and the full description text verbatim.
I appreciate your request, but I need to clarify what I'm looking at: The web page content you've provided is an **error page**, not an issue report. It displays: - An "Access Denied" message with error code b3728715388cb593 - A sad mascot image (Anubis) - Navigation and attribution information - A note that the site runs Anubis version "devel" There is **no issue title, status, author, dates, or description** present in this content to reproduce. This appears to be a standard 403/access denied error page from an Anubis-protected website, rather than a GitHub issue or similar documentation. If you're looking to report or reproduce an actual issue, could you provide the relevant issue tracker content instead?
Two, and I'd include both:
The commit (primary — stable, publicly readable, shows the diff and message):
https://github.com/illumos/illumos-gate/commit/68846fd00135fb0b10944e7806025cbefcfd6546
Use the full SHA rather than the abbreviated 68846fd00135 I had in the report — abbreviations can become ambiguous as a repo grows.
The bug (secondary, for the illumos-side discussion):
https://www.illumos.org/issues/4627
Caveat: I couldn't verify that page's contents. It returns 200 but serves an Anubis anti-bot challenge to automated fetchers. It should render fine in a browser, and 4627 is definitely the right number — it's in the commit message's first line — but worth a click before you paste it.
Let me fix the report to use the full SHA and add the tracker link:
Ran tools (1 Edit)
Edit /home/claude/golang_go_78438/report.md
/home/claude/golang_go_78438/report.md
**illumos fixed this exact defect** in commit [`68846fd00135`](https://github.com/illumos/illumos-gate/commit/68846fd00135fb0b10944e7806025cbefcfd6546) (2014-02-24), [illumos 4627](https://www.illumos.org/issues/4627), *"POLLHUP not generated for disconnected sockets"*. In
**illumos fixed this exact defect** in commit [`68846fd00135`](https://github.com/illumos/illumos-gate/commit/68846fd00135) (2014-02-24), *"4627 POLLHUP not generated for disconnected sockets"*. In
false
The file /home/claude/golang_go_78438/report.md has been updated successfully. (file state is current in your context — no need to Read it back)
The commit message references it as illumos 4627, "POLLHUP not generated for disconnected sockets" — a bare ID rather than a URL, which is the usual Go convention for third-party references. If you'd rather it carry the link, say so and I'll amend.
You Wed 26 Aug 2026 14:44:24
Are there other watchflakes filed issues that should be closed as duplicates of this one
Assistant Wed 26 Aug 2026 14:44:47
Good question — let me look. GitHub's issue search does index comments, which is where watchflakes records the builder names.
Ran tools (5 Bash)
Bash cd /tmp; q() { echo "--- $1"; curl -s -G "https://api.github.com/search/issues" --data-urlencode "q=$1" | python3 -c " ...
cd /tmp; q() { echo "--- $1"; curl -s -G "https://api.github.com/search/issues" --data-urlencode "q=$1" | python3 -c "
import json,sys
d=json.load(sys.stdin)
print('total',d.get('total_count'))
for i in d.get('items',[])[:20]:
print(' ', i['number'], i['state'], i['title'][:80])
"; sleep 2; }
q 'repo:golang/go "httptest.Server blocked in Close" in:comments'
q 'repo:golang/go "blocked in Close after 5 seconds" in:comments is:open'
q 'repo:golang/go label:Automation is:open solaris in:comments'Search issue comments for the hang signature
300000
--- repo:golang/go "httptest.Server blocked in Close" in:comments total 33 60310 open net/http: TestContentEncodingNoSniffing failures 67315 open net/http: TestTLSHandshakeTrace failures 60259 open net/http: frequent timeouts on `linux-amd64-wsl` with goroutines blocked in `che 19032 open net, net/http: unpredictable slowness of TestSendfile and TestOnlyWriteTimeout w 57612 open net/http: TestMaxBytesHandler failures 59447 open net/http: timeouts in TestResponseControllerSetPastReadDeadline 52896 open net/http: apparent deadlock in TestIdentityTransferEncoding_h1 via syscall.Open 78438 open net/http: TestTransportGzip/h1 failures 60359 open net/http: failures with `timeout waiting for SETTINGS frames` 75933 closed net/http/httputil: ReverseProxy stuck forever when handling request with expect: 65596 closed net/http: TestOmitHTTP2Vet failures 74187 open all: umbrella for flakes on openbsd/ppc64 49338 open all: plan9-arm builder failing frequently with a variety of errors 55164 open cmd/go: test failures due to network 60109 closed all: timeouts on `linux-s390x-ibm-race` 57369 closed net/http/cgi: TestCopyError failures due to unexpected child process 57599 closed net/http: TestClientInsecureTransport failures with `timeout waiting for logged 49336 closed net/http: failures ending in "Timeout." in testTransportAndServerSharedBodyRace 60413 open runtime: linux/s390x test timeout 57084 closed net/http: TestServerGracefulClose failures with `connection reset by peer` --- repo:golang/go "blocked in Close after 5 seconds" in:comments is:open total 13 60310 open net/http: TestContentEncodingNoSniffing failures 67315 open net/http: TestTLSHandshakeTrace failures 60259 open net/http: frequent timeouts on `linux-amd64-wsl` with goroutines blocked in `che 19032 open net, net/http: unpredictable slowness of TestSendfile and TestOnlyWriteTimeout w 57612 open net/http: TestMaxBytesHandler failures 59447 open net/http: timeouts in TestResponseControllerSetPastReadDeadline 52896 open net/http: apparent deadlock in TestIdentityTransferEncoding_h1 via syscall.Open 78438 open net/http: TestTransportGzip/h1 failures 60359 open net/http: failures with `timeout waiting for SETTINGS frames` 74187 open all: umbrella for flakes on openbsd/ppc64 49338 open all: plan9-arm builder failing frequently with a variety of errors 55164 open cmd/go: test failures due to network 60413 open runtime: linux/s390x test timeout --- repo:golang/go label:Automation is:open solaris in:comments total 6 78438 open net/http: TestTransportGzip/h1 failures 81106 open net/http: TestTransportBodyReadError/h3 failures 81075 open net: TestLookupCNAME failures 74941 open net/http: Test304Responses failures 73487 open runtime:cpu2: TestAbortInCgo failures 63444 open all: test failures with `EBADF` from fork/exec on NetBSD Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; python3 - <<'PY' ...
cd /tmp; python3 - <<'PY'
import json,urllib.request,re,time
cands=[60310,67315,57612,59447,52896,60359,19032,74941,81106]
def get(u):
r=urllib.request.Request(u, headers={'User-Agent':'triage'})
return json.load(urllib.request.urlopen(r))
for n in cands:
try:
body=get(f"https://api.github.com/repos/golang/go/issues/{n}")['body'] or ''
cs=get(f"https://api.github.com/repos/golang/go/issues/{n}/comments?per_page=100")
except Exception as e:
print(n,'ERR',e); continue
text=body+"\n".join(c['body'] for c in cs)
sol=[l for l in text.split('\n') if 'solaris' in l.lower() and ('<summary>' in l or 'gotip' in l or 'go1.' in l)]
if not sol:
print(f"{n}: no solaris failures")
continue
hasclose='blocked in Close' in text
haswrite='FD).Write' in text or 'FD).Write' in text
print(f"{n}: {len(sol)} solaris entries | 'blocked in Close' in issue: {hasclose} | 'FD).Write': {haswrite}")
for l in sol[:4]:
m=re.search(r'>([^<]*solaris[^<]*)<', l) or re.search(r'(\S*solaris\S*)', l)
print(' ', (m.group(1) if m else l)[:130])
time.sleep(1)
PYCheck candidate issues for solaris failures
300000
60310: no solaris failures
67315: no solaris failures
57612: no solaris failures
59447: no solaris failures
52896: no solaris failures
60359: no solaris failures
19032: no solaris failures
74941: 2 solaris entries | 'blocked in Close' in issue: False | 'FD).Write': False
2025-11-04 00:37 gotip-solaris-amd64 go@e2c6a202 net/http.Test304Responses [SKIP] (
2026-08-05 19:18 go1.25-solaris-amd64 release-branch.go1.25@bff0502b net/http.Test304Responses [SKIP] (
81106: 1 solaris entries | 'blocked in Close' in issue: False | 'FD).Write': False
2026-08-25 14:11 gotip-solaris-amd64 go@36042fe9 net/http.TestTransportBodyReadError/h3 (
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; python3 - <<'PY' ...
cd /tmp; python3 - <<'PY'
import json,urllib.request,re
def get(u):
r=urllib.request.Request(u, headers={'User-Agent':'triage'})
return json.load(urllib.request.urlopen(r))
for n in (74941,81106):
cs=get(f"https://api.github.com/repos/golang/go/issues/{n}/comments?per_page=100")
body=get(f"https://api.github.com/repos/golang/go/issues/{n}")
print('==',n,body['title'])
text=(body['body'] or '')+"\n".join(c['body'] for c in cs)
for l in text.split('\n'):
if 'solaris' in l.lower() and 'summary' in l:
m=re.search(r'https://ci\.chromium\.org/b/(\d+)', l)
print(' ', l[:120].replace('<details><summary>',''), '| build', m.group(1) if m else '?')
PYExtract solaris build IDs from candidate issues
180000
== 74941 net/http: Test304Responses failures
2025-11-04 00:37 gotip-solaris-amd64 go@e2c6a202 net/http.Test304Responses [SKIP] (<a href="https://ci | build 8699173838737036833
2026-08-05 19:18 go1.25-solaris-amd64 release-branch.go1.25@bff0502b net/http.Test304Responses [SKIP] | build 8674231095210346561
== 81106 net/http: TestTransportBodyReadError/h3 failures
2026-08-25 14:11 gotip-solaris-amd64 go@36042fe9 net/http.TestTransportBodyReadError/h3 (<a href="http | build 8672483468781074337
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; for id in 8699173838737036833 8674231095210346561; do ...
cd /tmp; for id in 8699173838737036833 8674231095210346561; do
( curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/$id/+/u/step/11/log/3?format=raw" | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('d_$id.txt','w').write(''.join(out))
" ) &
done; wait; ls -l d_*.txtFetch logs for Test304Responses solaris failures
600000
[1]- Done ( curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/$id/+/u/step/11/log/3?format=raw" | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('d_$id.txt','w').write(''.join(out))
" )
[2]+ Done ( curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/$id/+/u/step/11/log/3?format=raw" | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
open('d_$id.txt','w').write(''.join(out))
" )
-rw-rw-r-- 1 claude claude 427905 Aug 26 14:45 d_8674231095210346561.txt
-rw-rw-r-- 1 claude claude 451936 Aug 26 14:45 d_8699173838737036833.txt
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; for f in d_*.txt; do echo "===== $f"; grep -n "blocked in Close" $f | head -2; grep -n "running tests:" -A5 $f ...
cd /tmp; for f in d_*.txt; do echo "===== $f"; grep -n "blocked in Close" $f | head -2; grep -n "running tests:" -A5 $f | head -8; python3 - "$f" <<'EOF'
import re,sys
s=open(sys.argv[1]).read()
i=s.find('panic: test timed out')
if i<0: print(' no timeout panic'); raise SystemExit
for b in re.split(r'\n\n', s[i:]):
if 'FD).Write' in b or 'httptest.(*Server).Close' in b:
for l in b.split('\n'):
if any(k in l for k in ('FD).Write','chunkWriter','net/http_test.','httptest.(*Server).Close','goroutine ')):
print(' ', l.strip()[:110])
print(' --')
EOF
doneCheck signature of Test304Responses solaris hangs
===== d_8674231095210346561.txt
1068: running tests:
1069- TestTransportAndServerSharedBodyRace (8m59s)
1070- TestTransportAndServerSharedBodyRace/h1 (8m59s)
1071-
1072-goroutine 1088 gp=0xc000283340 m=9 mp=0xc000580008 [running]:
1073-panic({0x91c920?, 0xc0002e0520?})
goroutine 842 gp=0xc000405180 m=nil [IO wait, 8 minutes]:
internal/poll.(*FD).Write(0xc000360180, {0xc0003f4ff8, 0x7008, 0x7008})
net.(*netFD).Write(0xc000360180, {0xc0003f4ff8?, 0x1?, 0x0?})
net/http.(*chunkWriter).Write(0xc0003ae040, {0xc0003f4000, 0x8000, 0x8000})
net/http_test.testTransportAndServerSharedBodyRace.func1.1({0xab4dd0, 0xc0003ae000}, 0xc0003a8000)
created by net/http.(*Server).Serve in goroutine 1288
--
===== d_8699173838737036833.txt
1064: running tests:
1065- TestTransportAndServerSharedBodyRace (8m59s)
1066- TestTransportAndServerSharedBodyRace/h1 (8m59s)
1067-
1068-goroutine 1115 gp=0x33ef787cc380 m=0 mp=0xe7dc20 [running]:
1069-panic({0x932080?, 0x33ef78028b10?})
goroutine 1262 gp=0x33ef781a1500 m=nil [IO wait, 8 minutes]:
internal/poll.(*FD).Write(0x33ef78032a00, {0x33ef784bcff8, 0x7008, 0x7008})
net.(*netFD).Write(0x33ef78032a00, {0x33ef784bcff8?, 0x1?, 0x0?})
net/http.(*chunkWriter).Write(0x33ef78238e50, {0x33ef784bc000, 0x8000, 0x8000})
net/http_test.testTransportAndServerSharedBodyRace.func1.1({0xad27d0, 0x33ef78238e10}, 0x33ef78244a00)
created by net/http.(*Server).Serve in goroutine 1302
--
Shell cwd was reset to /home/claude/golang_go_78438Found something — those are the same bug under a different test name. Let me map the full extent.
Ran tools (3 Bash)
Bash cd /tmp; python3 - <<'PY' ...
cd /tmp; python3 - <<'PY'
import json,re
d=json.load(open('c60152.json'))
print('comments in #60152:', len(d))
seen=set()
for c in d:
for l in c['body'].split('\n'):
if '<summary>' in l and 'solaris' in l.lower():
m=re.search(r'https://ci\.chromium\.org/b/(\d+)', l)
txt=re.sub(r'<[^>]*>','',l).strip()
key=txt[:200]
if key in seen: continue
seen.add(key)
if 'net/http' in txt or 'net.' in txt:
print(' ', txt[:150], '| build', m.group(1) if m else '?')
PY
echo "=== open issues for the shared-body-race test ==="
curl -s -G "https://api.github.com/search/issues" --data-urlencode "q=repo:golang/go TransportAndServerSharedBodyRace" | python3 -c "
import json,sys
d=json.load(sys.stdin)
for i in d.get('items',[]): print(' ', i['number'], i['state'], i['title'][:90])"Scan solaris umbrella for net/http timeouts
300000
comments in #60152: 56 2023-08-20 05:56 solaris-amd64-oraclerel go@a2ec545e net/http (log) | build ? 2023-11-08 17:54 solaris-amd64-oraclerel go@ba0dd3b3 net/http.TestTransportAndServerSharedBodyRace (log) | build ? 2023-12-13 00:22 solaris-amd64-oraclerel go@400e24a8 net/http (log) | build ? 2025-02-21 17:17 gotip-solaris-amd64 go@d93f6df0 net/http.Test304Responses [ABORT] (log) | build 8722304712449172257 === open issues for the shared-body-race test === Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; python3 - <<'PY' ...
cd /tmp; python3 - <<'PY'
import json,re
d=json.load(open('c60152.json'))
for c in d:
for l in c['body'].split('\n'):
if '<summary>' in l and 'solaris' in l.lower() and ('net/http' in l or 'net.' in l):
print(l[:400]); print()
PYGet raw links for solaris net/http entries
<details><summary>2023-08-20 05:56 solaris-amd64-oraclerel go@a2ec545e net/http (<a href="https://build.golang.org/log/94fd1223950a4d9a14307c4052d6009997dbd3c8">log</a>)</summary> <details><summary>2023-11-08 17:54 solaris-amd64-oraclerel go@ba0dd3b3 net/http.TestTransportAndServerSharedBodyRace (<a href="https://build.golang.org/log/eef4abf515c9d219c7fc309ec6a2a464c5510af5">log</a>)</summary> <details><summary>2023-12-13 00:22 solaris-amd64-oraclerel go@400e24a8 net/http (<a href="https://build.golang.org/log/4b0843f00366705aef3feee76d784e24a20890dc">log</a>)</summary> <details><summary>2025-02-21 17:17 gotip-solaris-amd64 go@d93f6df0 net/http.Test304Responses [ABORT] (<a href="https://ci.chromium.org/b/8722304712449172257">log</a>)</summary> Shell cwd was reset to /home/claude/golang_go_78438
Bash cd /tmp; for u in eef4abf515c9d219c7fc309ec6a2a464c5510af5 4b0843f00366705aef3feee76d784e24a20890dc 94fd1223950a4d9a1430...
cd /tmp; for u in eef4abf515c9d219c7fc309ec6a2a464c5510af5 4b0843f00366705aef3feee76d784e24a20890dc 94fd1223950a4d9a14307c4052d6009997dbd3c8; do
echo "===== $u"; curl -sL "https://build.golang.org/log/$u" -o old_$u.txt; ls -l old_$u.txt | awk '{print $5" bytes"}'
grep -n "running tests:" -A4 old_$u.txt | head -7
grep -n "FD).Write" old_$u.txt | head -3
grep -n "blocked in Close" old_$u.txt | head -2
doneCheck 2023 solaris net/http hang signatures
300000
===== eef4abf515c9d219c7fc309ec6a2a464c5510af5
389681 bytes
174:running tests:
175- TestTransportAndServerSharedBodyRace (3m0s)
176- TestTransportAndServerSharedBodyRace/h1 (3m0s)
177-
178-goroutine 969 [running]:
3395:internal/poll.(*FD).Write(0xc0001f2800, {0xc000336ff8, 0x7008, 0x7008})
3397:net.(*netFD).Write(0xc0001f2800, {0xc000336ff8?, 0xc000336000?, 0x8000?})
===== 4b0843f00366705aef3feee76d784e24a20890dc
166690 bytes
176:running tests:
177- TestTransportDecrementConnWhenIdleConnRemoved/h2 (2m59s)
178-
179-goroutine 16229 gp=0xc000c31dc0 m=29 mp=0xc00055e408 [running]:
180-panic({0x871d40?, 0xc004c26f20?})
===== 94fd1223950a4d9a14307c4052d6009997dbd3c8
29948 bytes
171:running tests:
172- TestServerGracefulClose (3m0s)
173- TestServerGracefulClose/h1 (2m59s)
174-
175-goroutine 16796 [running]:
470:internal/poll.(*FD).Write(0xc000565000, {0xc0035e2000, 0x50003b, 0x526000})
472:net.(*netFD).Write(0xc000565000, {0xc0035e2000?, 0xc000468270?, 0x16?})
Shell cwd was reset to /home/claude/golang_go_78438Important find — this predates the drain CL. Let me verify and test that directly.
Ran tools (1 Bash)
Bash cd /tmp; sed -n '3390,3410p' old_eef4abf515c9d219c7fc309ec6a2a464c5510af5.txt; echo "=== 2023-08 GracefulClose ==="; sed...
cd /tmp; sed -n '3390,3410p' old_eef4abf515c9d219c7fc309ec6a2a464c5510af5.txt; echo "=== 2023-08 GracefulClose ==="; sed -n '465,490p' old_94fd1223950a4d9a14307c4052d6009997dbd3c8.txt; echo "=== gomote ==="; gomote list 2>&1 | grep solaris
Inspect 2023 stack frames
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/runtime/netpoll.go:345 +0x90 fp=0xc0000d13a8 sp=0xc0000d1388 pc=0x467ef0
internal/poll.(*pollDesc).wait(0xc0001f2800?, 0xc00033cff8?, 0x0)
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/internal/poll/fd_poll_runtime.go:84 +0x27 fp=0xc0000d13d0 sp=0xc0000d13a8 pc=0x49fee7
internal/poll.(*pollDesc).waitWrite(...)
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/internal/poll/fd_poll_runtime.go:93
internal/poll.(*FD).Write(0xc0001f2800, {0xc000336ff8, 0x7008, 0x7008})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/internal/poll/fd_unix.go:388 +0x2d9 fp=0xc0000d1480 sp=0xc0000d13d0 pc=0x4a1259
net.(*netFD).Write(0xc0001f2800, {0xc000336ff8?, 0xc000336000?, 0x8000?})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/net/fd_posix.go:96 +0x25 fp=0xc0000d14c8 sp=0xc0000d1480 pc=0x5eb4e5
net.(*conn).Write(0xc00014a048, {0xc000336ff8?, 0xc0000d1558?, 0xc0000be1b0?})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/net/net.go:191 +0x45 fp=0xc0000d1510 sp=0xc0000d14c8 pc=0x5f5fc5
net.(*TCPConn).Write(0x505815?, {0xc000336ff8?, 0xc000199000?, 0xc0000d1580?})
<autogenerated>:1 +0x25 fp=0xc0000d1540 sp=0xc0000d1510 pc=0x6024e5
net/http.checkConnErrorWriter.Write({0xc00019c000?}, {0xc000336ff8?, 0x4?, 0xc0000d1640?})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/net/http/server.go:3792 +0x2d fp=0xc0000d1590 sp=0xc0000d1540 pc=0x6fcf4d
bufio.(*Writer).Write(0xc00019c000, {0xc000336000?, 0x8f5bde?, 0x4?})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/bufio/bufio.go:679 +0xea fp=0xc0000d15f0 sp=0xc0000d1590 pc=0x505a4a
net/http.(*chunkWriter).Write(0xc000330040, {0xc000336000, 0x8000, 0x8000})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/net/http/server.go:387 +0x106 fp=0xc0000d1660 sp=0xc0000d15f0 pc=0x6edf66
bufio.(*Writer).Write(0xc000318080, {0xc000336000?, 0x0?, 0xc0000d1710?})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/bufio/bufio.go:679 +0xea fp=0xc0000d16c0 sp=0xc0000d1660 pc=0x505a4a
=== 2023-08 GracefulClose ===
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/runtime/netpoll.go:345 +0x90 fp=0xc000f4fe00 sp=0xc000f4fde0 pc=0x467e70
internal/poll.(*pollDesc).wait(0xa?, 0xc003629e50?, 0x0)
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/internal/poll/fd_poll_runtime.go:84 +0x27 fp=0xc000f4fe28 sp=0xc000f4fe00 pc=0x49f967
internal/poll.(*pollDesc).waitWrite(...)
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/internal/poll/fd_poll_runtime.go:93
internal/poll.(*FD).Write(0xc000565000, {0xc0035e2000, 0x50003b, 0x526000})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/internal/poll/fd_unix.go:388 +0x2c8 fp=0xc000f4fed8 sp=0xc000f4fe28 pc=0x4a0cc8
net.(*netFD).Write(0xc000565000, {0xc0035e2000?, 0xc000468270?, 0x16?})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/net/fd_posix.go:96 +0x25 fp=0xc000f4ff20 sp=0xc000f4fed8 pc=0x5e6ee5
net.(*conn).Write(0xc00006c288, {0xc0035e2000?, 0xc0007cc198?, 0xc002632600?})
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/net/net.go:191 +0x45 fp=0xc000f4ff68 sp=0xc000f4ff20 pc=0x5f1805
net.(*TCPConn).Write(0xc000015d05?, {0xc0035e2000?, 0x7cd7a0?, 0xc000015d30?})
<autogenerated>:1 +0x25 fp=0xc000f4ff98 sp=0xc000f4ff68 pc=0x5fe925
net/http_test.testServerGracefulClose.func2()
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/net/http/serve_test.go:3097 +0x31 fp=0xc000f4ffe0 sp=0xc000f4ff98 pc=0x796fd1
runtime.goexit()
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/runtime/asm_amd64.s:1650 +0x1 fp=0xc000f4ffe8 sp=0xc000f4ffe0 pc=0x46e421
created by net/http_test.testServerGracefulClose in goroutine 1898
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/net/http/serve_test.go:3096 +0x29f
goroutine 5683 [IO wait, 2 minutes]:
runtime.gopark(0xc0009d8c40?, 0x435076?, 0x88?, 0x8c?, 0x4a285d?)
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/runtime/proc.go:398 +0xce fp=0xc0009d8c18 sp=0xc0009d8bf8 pc=0x43bfee
runtime.netpollblock(0x7fffbe21d950?, 0x72?, 0x0?)
/tmp/workdir-host-solaris-oracle-amd64-oraclerel/go/src/runtime/netpoll.go:573 +0xf7 fp=0xc0009d8c50 sp=0xc0009d8c18 pc=0x434777
internal/poll.runtime_pollWait(0x7fffbe21d950, 0x72)
=== gomote ===The gomote expired. Let me spin up a new one and test the CL 737720 claim directly, since the 2023 evidence casts doubt on it.
Ran tools (8 Bash)
Bash cd /home/claude/golang_go_78438/go && (gomote create gotip-solaris-amd64 > /tmp/gm.log 2>&1 &) ; git checkout -q master ...
cd /home/claude/golang_go_78438/go && (gomote create gotip-solaris-amd64 > /tmp/gm.log 2>&1 &) ; git checkout -q master && python3 - <<'PY' p='src/net/http/transport.go' s=open(p).read() old='tryDrain := !bodyEOF && resp.ContentLength <= maxPostCloseReadBytes' new='tryDrain := false && !bodyEOF && resp.ContentLength <= maxPostCloseReadBytes' assert old in s open(p,'w').write(s.replace(old,new,1)) PY SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad cd src && GOOS=solaris GOARCH=amd64 ../bin/go test -c -o $SP/nethttp.nodrain.test net/http 2>&1|tail -3 cd .. && git checkout -q src/net/http/transport.go && git checkout -q fix-78438-write-hang && git status --porcelain; ls -l $SP/nethttp.nodrain.test
Build no-drain baseline binary
300000
-rwxrwxr-x 1 claude claude 17900501 Aug 26 14:47 /tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad/nethttp.nodrain.test
Bash cd /tmp; python3 - <<'PY' ...
cd /tmp; python3 - <<'PY'
import json,urllib.request,re
def get(u):
r=urllib.request.Request(u, headers={'User-Agent':'triage'})
return json.load(urllib.request.urlopen(r))
for n in (74941,81106):
b=get(f"https://api.github.com/repos/golang/go/issues/{n}")
cs=get(f"https://api.github.com/repos/golang/go/issues/{n}/comments?per_page=100")
text=(b['body'] or '')+"\n".join(c['body'] for c in cs)
ents=[re.sub(r'<[^>]*>','',l).strip() for l in text.split('\n') if '<summary>' in l]
print(f"== #{n} {b['title']} ({len(ents)} failure entries)")
for e in ents: print(' ', e[:120])
print()
PYList all failure entries for candidate issues
180000
== #74941 net/http: Test304Responses failures (6 failure entries)
2025-08-08 02:04 gotip-darwin-amd64-race go@73ff6d14 net/http.Test304Responses [SKIP] (log)
2025-08-12 16:00 gotip-darwin-amd64-race go@d0a64f79 net/http.Test304Responses [SKIP] (log)
2025-11-04 00:37 gotip-solaris-amd64 go@e2c6a202 net/http.Test304Responses [SKIP] (log)
2026-03-31 17:58 gotip-linux-ppc64le_power10 go@f665ff8b net/http.Test304Responses [SKIP] (log)
2026-04-23 15:12 gotip-linux-ppc64_power8 go@13cab13f net/http.Test304Responses [SKIP] (log)
2026-08-05 19:18 go1.25-solaris-amd64 release-branch.go1.25@bff0502b net/http.Test304Responses [SKIP] (log)
== #81106 net/http: TestTransportBodyReadError/h3 failures (2 failure entries)
2026-08-25 14:11 gotip-solaris-amd64 go@36042fe9 net/http.TestTransportBodyReadError/h3 (log)
2026-08-25 17:11 gotip-linux-mips64le go@b93d8195 net/http.TestTransportBodyReadError/h3 (log)
Shell cwd was reset to /home/claude/golang_go_78438Bash cd /tmp; curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672483468781074337/+/u/step/11/log/...
cd /tmp; curl -sL "https://logs.chromium.org/logs/golang/buildbucket/cr-buildbucket/8672483468781074337/+/u/step/11/log/3?format=raw" 2>/dev/null | python3 -c "
import json,sys
out=[]
for line in sys.stdin:
try: d=json.loads(line)
except: continue
if d.get('Package')!='net/http': continue
t=d.get('Output','')
if t: out.append(t)
s=''.join(out)
open('h3.txt','w').write(s)
import re
i=s.find('TestTransportBodyReadError')
print('has timeout panic:', 'panic: test timed out' in s)
print('has blocked in Close:', 'blocked in Close' in s)
print('has FD).Write:', 'FD).Write' in s)
for l in s.split('\n'):
if 'FAIL: TestTransportBodyReadError' in l or 'transport_test.go' in l and 'Error' in l:
print(' ', l.strip()[:160])
"; tail -1 /tmp/gm.logCheck the h3 solaris failure and gomote status
600000
has timeout panic: False has blocked in Close: False has FD).Write: False --- FAIL: TestTransportBodyReadError/h3 (0.09s) --- FAIL: TestTransportBodyReadError (0.00s) # still creating gotip-solaris-amd64 (1) after 32s; 0 requests ahead of you Shell cwd was reset to /home/claude/golang_go_78438
Bash for i in 1 2 3 4 5 6 7 8 9 10; do G=$(gomote list 2>/dev/null | grep solaris | awk '{print $1}'); if [ -n "$G" ]; then e...
for i in 1 2 3 4 5 6 7 8 9 10; do G=$(gomote list 2>/dev/null | grep solaris | awk '{print $1}'); if [ -n "$G" ]; then echo "ready: $G"; break; fi; sleep 20; done; tail -1 /tmp/gm.logWait for gomote
600000
ready: rsc-gotip-solaris-amd64-0 rsc-gotip-solaris-amd64-0
Bash SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-am...
SP=/tmp/claude-1003/-home-claude-golang-go-78438/399e501a-6ca0-43ff-9dd2-f1fc86d93995/scratchpad; G=rsc-gotip-solaris-amd64-0 gomote put $G $SP/stress.sh stress.sh 2>&1|tail -1 gomote put $G $SP/nethttp.nodrain.test nethttp.nodrain.test 2>&1|tail -1 gomote put $G $SP/nethttp.test nethttp.test 2>&1|tail -1 gomote run $G /bin/sh -c 'chmod +x stress.sh nethttp.nodrain.test nethttp.test; ./stress.sh nethttp.test base; ./stress.sh nethttp.nodrain.test nodrain; ./stress.sh nethttp.nodrain.test nodrain2' 2>&1 | tail -6
Test whether the drain is required for the hang
1500000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2277259389/rsc-gotip-solaris-amd64-0.stdout"... RESULT base: 1 of 8 wedged RESULT nodrain: 0 of 8 wedged RESULT nodrain2: 0 of 8 wedged # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote2277259389/rsc-gotip-solaris-amd64-0.stdout".
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c './stress.sh nethttp.test base3; ./stress.sh nethttp.nodrain.test ...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c './stress.sh nethttp.test base3; ./stress.sh nethttp.nodrain.test nodrain3; ./stress.sh nethttp.test base4; ./stress.sh nethttp.nodrain.test nodrain4' 2>&1 | tail -7
More interleaved rounds for signal
1800000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3722436115/rsc-gotip-solaris-amd64-0.stdout"... RESULT base3: 2 of 8 wedged RESULT nodrain3: 0 of 8 wedged RESULT base4: 5 of 8 wedged RESULT nodrain4: 1 of 8 wedged # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3722436115/rsc-gotip-solaris-amd64-0.stdout".
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'for i in 0 1 2 3 4 5 6 7; do if grep -q "test timed out" t.nodrai...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'for i in 0 1 2 3 4 5 6 7; do if grep -q "test timed out" t.nodrain4.$i.out; then echo "=== t.nodrain4.$i.out"; grep -n "blocked in Close" t.nodrain4.$i.out | head -1; grep -n "running tests:" -A2 t.nodrain4.$i.out | head -3; grep -n "FD).Write\|testTransportGzip" t.nodrain4.$i.out | head -3; fi; done' 2>&1 | tail -12
Confirm no-drain wedge has same signature
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote847287273/rsc-gotip-solaris-amd64-0.stdout"... === t.nodrain4.3.out 1:2026/08/26 21:01:10 httptest.Server blocked in Close after 5 seconds, waiting for connections: grep: can't open -A2 t.nodrain4.3.out:4: running tests: # Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote847287273/rsc-gotip-solaris-amd64-0.stdout".
Bash G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'sed -n "1,8p" t.nodrain4.3.out; echo ...; grep -n "FD).Write" t.n...
G=rsc-gotip-solaris-amd64-0; gomote run $G /bin/sh -c 'sed -n "1,8p" t.nodrain4.3.out; echo ...; grep -n "FD).Write" t.nodrain4.3.out | head -2; grep -n "testTransportGzip" t.nodrain4.3.out | head -2' 2>&1 | tail -14
Show no-drain wedge details
300000
# Streaming results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3871399031/rsc-gotip-solaris-amd64-0.stdout"...
2026/08/26 21:01:10 httptest.Server blocked in Close after 5 seconds, waiting for connections:
*net.TCPConn 0x12bf9baaa060 127.0.0.1:54821 in state active
panic: test timed out after 3m0s
running tests:
TestTransportGzip (2m33s)
TestTransportGzip/h1 (2m33s)
goroutine 22029 [running]:
...
71:internal/poll.(*FD).Write(0x12bf9baa2100, {0x12bf9c47ef97, 0xff10d, 0x11f069})
73:net.(*netFD).Write(0x12bf9baa2100, {0x12bf9c47ef97?, 0x0?, 0x0?})
95:net/http_test.testTransportGzip.func1({0xf52fc0, 0x12bf9b99e1e0}, 0x12bf9bcd4140)
# Wrote results from "rsc-gotip-solaris-amd64-0" to "/tmp/gomote3871399031/rsc-gotip-solaris-amd64-0.stdout".Same signature without the drain. Correcting the report and adding the duplicate analysis:
Ran tools (2 Edit, 1 Bash, 1 SendUserFile)
Edit /home/claude/golang_go_78438/report.md
/home/claude/golang_go_78438/report.md
This is long-standing, not new. The same signature — a `net/http` goroutine parked forever in `internal/poll.(*FD).Write` on solaris — appears in #60152 as far back as 2023: - 2023-08-20 [`TestServerGracefulClose`](https://build.golang.org/log/94fd1223950a4d9a14307c4052d6009997dbd3c8) - 2023-11-08 [`TestTransportAndServerSharedBodyRace`](https://build.golang.org/log/eef4abf515c9d219c7fc309ec6a2a464c5510af5) CL 737720 (`11d5284363`, 2026-01-21), which made `Response.Body.Close` drain up to 256 KiB / 50 ms before closing, makes it far more likely without being required. The extra wake/write/re-arm cycles raise the odds that the reset lands while only `POLLIN` is armed, so the delivered event is `POLLIN`-only. Measured by stressing `TestTransportGzip` with the drain compiled out (`tryDrain := false`), interleaved with unmodified runs on the same machine: ``` as shipped 1, 2, 5, 2, 3, 4 wedged of 8 (17/48) drain disabled 0, 0, 0, 1 wedged of 8 ( 1/32) ``` The single `drain disabled` wedge has the identical signature, so the drain is an accelerant, roughly tenfold, not the cause. That fits both the 2023 failures and this test's first appearance in March.
The window widened noticeably after CL 737720 (`11d5284363`, 2026-01-21), which made `Response.Body.Close` drain up to 256 KiB / 50 ms before closing. The extra wake/write/re-arm cycles make it much more likely the reset lands while only `POLLIN` is armed, so the delivered event is `POLLIN`-only. That CL widened an existing kernel bug; it did not introduce one. It matches the first failure appearing in March.
false
The file /home/claude/golang_go_78438/report.md has been updated successfully. (file state is current in your context — no need to Read it back)
Edit /home/claude/golang_go_78438/report.md
/home/claude/golang_go_78438/report.md
A runtime-level fix was considered and rejected. In the wedged state only `POLLOUT` is armed, so waking writers on `POLLIN` events does nothing; the only mask that fires is `POLLIN`, and arming `POLLIN` alongside `POLLOUT` for writers would spin, because `POLLIN` on a socket with unread data the application is not reading is level-triggered and permanent. The real fix belongs in Solaris. Note the mitigation covers `net/http` servers only. The 2023-08-20 failure above is a plain `net.TCPConn.Write` in test code, with no `net/http` server involved; that shape stays exposed until Solaris is fixed. ## Other issues affected No other watchflakes issue is a whole-issue duplicate, but some of their failures belong here and should stop when this lands: - **#74941 (`Test304Responses`)** — all six entries are `[SKIP]`, meaning `Test304Responses` was merely paused when the binary died; it is not the failing test in any of them. Both solaris entries (2025-11-04, 2026-08-05) are this bug, with `TestTransportAndServerSharedBodyRace/h1` wedged in `internal/poll.(*FD).Write`. The other four are darwin-amd64-race ×2 and linux-ppc64le/ppc64, presumably unrelated. The issue is misfiled rather than duplicate: it is not really about `Test304Responses` at all, and would be worth re-triaging from its logs. - **#60152 (solaris timeouts umbrella)** — holds at least three failures of this exact class (2023-08-20, 2023-11-08, and 2025-02-21 `Test304Responses [ABORT]`). Stays open; it is broader than this. - **#78576** — the one linux-ppc64le failure on this issue belongs there, not here. Checked and *not* related: **#81106** (`TestTransportBodyReadError/h3`), whose solaris entry is a genuine assertion failure in 0.09s, no hang.
A runtime-level fix was considered and rejected. In the wedged state only `POLLOUT` is armed, so waking writers on `POLLIN` events does nothing; the only mask that fires is `POLLIN`, and arming `POLLIN` alongside `POLLOUT` for writers would spin, because `POLLIN` on a socket with unread data the application is not reading is level-triggered and permanent. The real fix belongs in Solaris.
false
The file /home/claude/golang_go_78438/report.md has been updated successfully. (file state is current in your context — no need to Read it back)
Bash cd /home/claude/golang_go_78438/go && git log --oneline -1 && git status --porcelain && echo "--- report ---" && wc -l ....
cd /home/claude/golang_go_78438/go && git log --oneline -1 && git status --porcelain && echo "--- report ---" && wc -l ../report.md
Confirm state
3f3c7a57ca net/http: abort in-flight response writes when the connection dies --- report --- 266 ../report.md
SendUserFile
Updated: corrected the CL 737720 claim (measured, not inferred) and added the "Other issues affected" section.
[ "/home/claude/golang_go_78438/report.md" ]
normal
1 file delivered to user. /home/claude/golang_go_78438/report.md → file_uuid: 03095fbb-0969-4021-8d65-f937f2b63829
No clean duplicates to close, but one issue is misfiled and two others have failures that belong here.
#74941 net/http: Test304Responses failures — misfiled, not duplicate
All six of its entries are [SKIP], which means Test304Responses was merely paused when the binary died — it is not the failing test in any of them. Both solaris entries are this bug, with a different test wedged:
2025-11-04 gotip-solaris-amd64 → TestTransportAndServerSharedBodyRace/h1
2026-08-05 go1.25-solaris-amd64 → TestTransportAndServerSharedBodyRace/h1
both parked in internal/poll.(*FD).Write inside the handler. The other four are darwin-amd64-race ×2 and linux-ppc64le/ppc64, presumably unrelated. So it's not a duplicate of #78438 — it's an issue that isn't really about its own title, and deserves re-triage from its logs rather than closure.
#60152 (solaris umbrella) — holds three more of these
Stays open, it's broader, but worth noting it already contained this bug: 2023-08-20 TestServerGracefulClose, 2023-11-08 TestTransportAndServerSharedBodyRace, 2025-02-21 Test304Responses [ABORT].
#81106 — checked, unrelated
Its solaris entry is a genuine assertion failure in 0.09s, no hang.
This changed a claim in the report
Those 2023 failures predate CL 737720 by two years, so I tested the drain claim instead of leaving it inferred. Building TestTransportGzip with tryDrain := false and interleaving with unmodified runs on the same machine:
| wedged of 8 | |
|---|---|
| as shipped | 1, 2, 5, 2, 3, 4 → 17/48 |
| drain disabled | 0, 0, 0, 1 → 1/32 |
The one drain disabled wedge has the identical signature. So the drain is an accelerant of roughly tenfold, not the cause — my earlier "widened the window" wording was directionally right but I'd been treating it as the trigger for this test's March appearance. The report now states the measured numbers and cites the 2023 failures as prior art.
One consequence worth flagging: the 2023-08-20 failure is a plain net.TCPConn.Write in test code with no net/http server involved. My committed fix does nothing for that shape — it's server-side only. That's now called out explicitly in the report's mitigation section, so nobody reads the fix as closing off the whole class.