fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22) (#24)

* fix(ci): bump opencode Azure apiVersion to support /responses endpoint (#22)

The CUA smoke probe was returning MODEL_CAPABLE=false because opencode's
@ai-sdk/azure provider got 'API version not supported' from Azure on
/openai/v1/responses with apiVersion=2024-08-01-preview. Split into two
envs: keep the CUA driver on 2024-08-01-preview (chat-completions only)
and bump the opencode-side provider config to 2025-04-01-preview, which
supports the new responses API.

Effect: send_message/multi_turn scenarios get included again in CUA smoke
when the probe succeeds.

* ci(cua): bound curl timeouts + diag dump on server-start hang (#22)

Step 11 'Start opencode server' has hung past the 45-min job timeout in
two consecutive runs (27195348071, 27198238465). Local boot of
opencode-ai 1.16.2 with the same heredoc config is healthy in 3s, so
something is wrong specifically on the GH-hosted runner — likely curl
post-loop waiting indefinitely on an unresponsive server.

Adds:
- set -x for command tracing
- --connect-timeout 2 -m 5 on every curl so hangs cannot exceed 5s
- HEALTHY flag + explicit exit 1 (drops the unbounded post-loop curl)
- Periodic dump every 10s: server log tail, ss listening sockets,
  process liveness — so we can see WHY the server isn't replying

Pure diagnostics; no behaviour change for the green path.

* fix(ci): use api-version=preview for /openai/v1/responses (#22)

Reproduced the probe failure locally against the same Azure resource:
all date-based api-versions (2024-08-01-preview, 2024-12-01-preview,
2025-01-01-preview, 2025-03-01-preview, 2025-04-01-preview) return:

    {"error":{"code":"BadRequest","message":"API version not supported"}}

Only api-version=preview and api-version=v1 succeed (200). This is the
new Azure OpenAI v1 responses-API style; date strings are reserved for
the legacy /openai/deployments/{model}/chat/completions endpoint.

@ai-sdk/azure 3.x already defaults apiVersion to "preview" (per the
type definition: "Custom api version to use. Defaults to `preview`."),
so this aligns the workflow with the SDK default. Probe should now
return MODEL_CAPABLE=true and the send_message scenario will run.

* test(cua): extend send_message and multi_turn waits to 30s

Assistant bubbles can take 15+ seconds to appear after send. Previous
5-second wait was too short and caused false failures even when API
calls succeeded. Re-check screenshots periodically up to 30s total.

* fix(cua): screen-relative send button threshold for #22

The send action's auto-locate filtered for y1 > 2200 and fell back to
hardcoded (996, 2358) — both assume a 1080x2400 panel. The CI emulator
(API 30 google_apis pixel profile) is 1080x1920, so:
  - the bottom_buttons filter never matched any clickable element
  - the fallback tap landed off-screen
  → 'ping' message never sent, scenario timed out with no bubbles.

Switch to a screen-relative threshold (bottom 25%) and a fallback that
uses get_screen_size() to land in the bottom-right corner regardless of
device resolution. This was masked until now because send_message was
gated by MODEL_CAPABLE=false in earlier CI runs.

Refs: #22

---------

Co-authored-by: dzianisv <dzianis.varabyou@gmail.com>
This commit is contained in:
Den
2026-06-13 23:34:48 -07:00
committed by GitHub
parent a9822048e4
commit 71a1feb232
2 changed files with 57 additions and 18 deletions

View File

@@ -18,7 +18,15 @@ jobs:
env:
AZURE_OPENAI_API_KEY: ${{ secrets.AZURE_OPENAI_API_KEY }}
AZURE_OPENAI_ENDPOINT: ${{ secrets.AZURE_OPENAI_ENDPOINT }}
# CUA driver uses chat-completions; 2024-08-01-preview is sufficient.
AZURE_OPENAI_API_VERSION: "2024-08-01-preview"
# opencode's @ai-sdk/azure provider hits the new /openai/v1/responses
# endpoint. That endpoint accepts ONLY api-version=preview or v1 — every
# date-based value (2024-*-preview, 2025-*-preview, 2025-04-01-preview)
# returns 400 "API version not supported" and forces MODEL_CAPABLE=false,
# which makes send_message/multi_turn scenarios get skipped (#22).
# `preview` is also the @ai-sdk/azure 3.x default.
OPENCODE_AZURE_API_VERSION: "preview"
# Android emulator reaches the runner host loopback via 10.0.2.2.
# A real `opencode serve` runs on the host (see steps below), making this a true E2E.
OPENCODE_URL: "http://10.0.2.2:4096"
@@ -92,7 +100,7 @@ jobs:
"options": {
"resourceName": "${RESOURCE_NAME}",
"apiKey": "{env:AZURE_OPENAI_API_KEY}",
"apiVersion": "${AZURE_OPENAI_API_VERSION}"
"apiVersion": "${OPENCODE_AZURE_API_VERSION}"
},
"models": {
"${AZURE_OPENAI_MODEL}": { "name": "Azure ${AZURE_OPENAI_MODEL}" }
@@ -107,18 +115,40 @@ jobs:
- name: Start opencode server on runner host
run: |
set -x
# Bind all interfaces so the emulator can reach it via 10.0.2.2.
nohup opencode serve --hostname 0.0.0.0 --port 4096 > /tmp/opencode-server.log 2>&1 &
SRV_PID=$!
echo "opencode pid=$SRV_PID"
echo "Waiting for opencode server /global/health ..."
HEALTHY=0
for i in $(seq 1 60); do
if curl -sf http://127.0.0.1:4096/global/health > /dev/null; then
echo "opencode server healthy after ${i}s"; break
if curl -sf --connect-timeout 2 -m 5 http://127.0.0.1:4096/global/health > /dev/null; then
echo "opencode server healthy after ${i}s"
HEALTHY=1
break
fi
# Periodic state dump every 10s while waiting
if [ $((i % 10)) -eq 0 ]; then
echo "--- @${i}s: server log so far ---"
tail -20 /tmp/opencode-server.log || true
echo "--- listening ports ---"
ss -tlnp 2>/dev/null | grep -E ':4096|opencode' || echo "no listener on 4096"
echo "--- pid alive? ---"
kill -0 $SRV_PID 2>/dev/null && echo "pid $SRV_PID alive" || echo "pid $SRV_PID DEAD"
fi
sleep 1
done
curl -sf http://127.0.0.1:4096/global/health || {
echo "::error::opencode server failed to become healthy"; cat /tmp/opencode-server.log; exit 1;
}
if [ "$HEALTHY" != "1" ]; then
echo "::error::opencode server failed to become healthy in 60s"
echo "--- final server log ---"
cat /tmp/opencode-server.log || true
echo "--- ss listing ---"
ss -tlnp 2>/dev/null || true
echo "--- ps tree ---"
ps -ef | grep -E 'opencode|node|npm' | head -20 || true
exit 1
fi
- name: Probe opencode model capability (can it reply?)
id: probe