Files
opencode-mobile/.github/workflows/cua-smoke.yml
engineer cfb0d9fe32 test(ci): widen cua-smoke gate to real core journey (connect→send→reply→list)
The CI opencode server had NO LLM provider configured — the server log only
showed "listening", never a model. opencode-ai (released npm pkg) does not read
AZURE_OPENAI_* for its own LLM; it needs an explicit provider in opencode.json +
a default `model`. So send_message/multi_turn could never pass and the gate was
stuck on --only-connect-scenario (UI journey minus the model reply).

- Wire opencode to the same Azure resource the CUA driver uses via a generated
  ~/.config/opencode/opencode.json (@ai-sdk/azure provider, resourceName derived
  from the endpoint secret at runtime, apiKey from env, default model azure/gpt-5.4).
- Add a deterministic REST probe step: create a session + send a prompt and check
  for an assistant reply BEFORE the ~30min emulator run, exporting MODEL_CAPABLE.
- Add --scenarios to android-cua-smoke.py to run an explicit named set.
- Emulator step now runs connect_and_verify_sessions + send_message +
  verify_session_list when MODEL_CAPABLE=true; falls back to the UI-only journey
  (connect + verify_session_list) otherwise, logging the environmental reason.
- Raise --max-steps to 40 so multiple scenarios fit.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-08 05:58:19 -07:00

192 lines
8.3 KiB
YAML

name: CUA Smoke Test
on:
workflow_dispatch:
push:
branches: [main]
tags: ["v*"]
paths:
- "scripts/android-cua-smoke.py"
- "src/**"
- "app/**"
- ".github/workflows/cua-smoke.yml"
jobs:
cua-test:
runs-on: ubuntu-latest
timeout-minutes: 45
env:
AZURE_OPENAI_API_KEY: ${{ secrets.AZURE_OPENAI_API_KEY }}
AZURE_OPENAI_ENDPOINT: ${{ secrets.AZURE_OPENAI_ENDPOINT }}
AZURE_OPENAI_API_VERSION: "2024-08-01-preview"
# Android emulator reaches the runner host loopback via 10.0.2.2.
# A real `opencode serve` runs on the host (see steps below), making this a true E2E.
OPENCODE_URL: "http://10.0.2.2:4096"
ARCHIVEBOX_URL: ${{ secrets.ARCHIVEBOX_URL }}
ARCHIVEBOX_API_KEY: ${{ secrets.ARCHIVEBOX_API_KEY }}
steps:
- uses: actions/checkout@v6
- uses: actions/setup-node@v6
with:
node-version: 20
cache: npm
- uses: actions/setup-java@v5
with:
distribution: temurin
java-version: 17
- name: Setup Android SDK
uses: android-actions/setup-android@v4
- name: Enable KVM
run: |
echo 'KERNEL=="kvm", GROUP="kvm", MODE="0666", OPTIONS+="static_node=kvm"' | sudo tee /etc/udev/rules.d/99-kvm4all.rules
sudo udevadm control --reload-rules
sudo udevadm trigger --name-match=kvm
- name: Cache Gradle
uses: actions/cache@v5
with:
path: |
~/.gradle/caches
~/.gradle/wrapper
android/.gradle
key: ${{ runner.os }}-gradle-${{ hashFiles('android/**/*.gradle*', 'android/gradle/wrapper/gradle-wrapper.properties') }}
restore-keys: |
${{ runner.os }}-gradle-
- name: Install dependencies & build APK
env:
SENTRY_DISABLE_AUTO_UPLOAD: "true"
run: |
npm install --legacy-peer-deps
npx expo prebuild --platform android --no-install
keytool -genkey -v -keystore android/app/debug.keystore -storepass android -alias androiddebugkey -keypass android -keyalg RSA -keysize 2048 -validity 10000 -dname "CN=Android Debug,O=Android,C=US"
cd android && ./gradlew assembleRelease
- name: Install Python deps
run: pip install openai
- name: Configure opencode Azure provider
env:
AZURE_OPENAI_MODEL: "gpt-5.4"
run: |
npm install -g opencode-ai
# opencode (released npm pkg) does NOT read AZURE_OPENAI_* for its own LLM.
# It needs an explicit provider in opencode.json + a default `model`.
# Wire the same Azure resource the CUA driver uses (@ai-sdk/azure).
# Extract the resource name from the endpoint secret at runtime
# (e.g. https://NAME.openai.azure.com -> NAME) so nothing secret is in source.
RESOURCE_NAME="$(printf '%s' "$AZURE_OPENAI_ENDPOINT" | sed -E 's#https?://([^.]+)\..*#\1#')"
echo "Derived Azure resource name: ${RESOURCE_NAME:-<empty>}"
mkdir -p "$HOME/.config/opencode"
cat > "$HOME/.config/opencode/opencode.json" <<EOF
{
"\$schema": "https://opencode.ai/config.json",
"provider": {
"azure": {
"npm": "@ai-sdk/azure",
"name": "Azure OpenAI",
"options": {
"resourceName": "${RESOURCE_NAME}",
"apiKey": "{env:AZURE_OPENAI_API_KEY}",
"apiVersion": "${AZURE_OPENAI_API_VERSION}"
},
"models": {
"${AZURE_OPENAI_MODEL}": { "name": "Azure ${AZURE_OPENAI_MODEL}" }
}
}
},
"model": "azure/${AZURE_OPENAI_MODEL}",
"small_model": "azure/${AZURE_OPENAI_MODEL}"
}
EOF
echo "opencode.json written:"; cat "$HOME/.config/opencode/opencode.json"
- name: Start opencode server on runner host
run: |
# Bind all interfaces so the emulator can reach it via 10.0.2.2.
nohup opencode serve --hostname 0.0.0.0 --port 4096 > /tmp/opencode-server.log 2>&1 &
echo "Waiting for opencode server /global/health ..."
for i in $(seq 1 60); do
if curl -sf http://127.0.0.1:4096/global/health > /dev/null; then
echo "opencode server healthy after ${i}s"; break
fi
sleep 1
done
curl -sf http://127.0.0.1:4096/global/health || {
echo "::error::opencode server failed to become healthy"; cat /tmp/opencode-server.log; exit 1;
}
- name: Probe opencode model capability (can it reply?)
id: probe
run: |
# Deterministic check that the server can actually produce an assistant
# reply BEFORE we spend ~30min driving the UI. Creates a session, sends a
# prompt via REST, and checks for an assistant message. Non-fatal: records
# MODEL_CAPABLE=true/false so the scenario set can be chosen accordingly.
set +e
SID=$(curl -sf -X POST http://127.0.0.1:4096/session -H 'content-type: application/json' -d '{}' | python3 -c "import sys,json;print(json.load(sys.stdin).get('id',''))" 2>/dev/null)
echo "session id: ${SID:-<none>}"
CAPABLE=false
if [ -n "$SID" ]; then
curl -sf -X POST "http://127.0.0.1:4096/session/$SID/message" \
-H 'content-type: application/json' \
-d '{"parts":[{"type":"text","text":"reply with the single word: pong"}]}' \
> /tmp/probe_reply.json 2>/tmp/probe_err.txt
echo "--- probe reply (truncated) ---"; head -c 2000 /tmp/probe_reply.json; echo
echo "--- probe err (truncated) ---"; head -c 1000 /tmp/probe_err.txt; echo
if grep -qi 'assistant' /tmp/probe_reply.json && grep -qi 'pong\|text' /tmp/probe_reply.json; then
CAPABLE=true
fi
fi
echo "MODEL_CAPABLE=$CAPABLE" >> "$GITHUB_OUTPUT"
echo "opencode model-capable in CI: $CAPABLE"
- name: Run CUA smoke test with emulator
uses: reactivecircus/android-emulator-runner@v2
env:
MODEL_CAPABLE: ${{ steps.probe.outputs.MODEL_CAPABLE }}
with:
api-level: 30
arch: x86_64
target: google_apis
emulator-options: -no-window -no-audio -no-boot-anim -gpu swiftshader_indirect -no-snapshot
# NOTE: android-emulator-runner runs this script with /usr/bin/sh (dash).
# Avoid bash-only constructs like multi-line `|| { ... }` brace groups.
script: |
# Non-fatal re-check; server health was already gated in the prior step.
curl -sf http://127.0.0.1:4096/global/health || echo "WARN: opencode health re-check failed (started in prior step)"
adb install android/app/build/outputs/apk/release/app-release.apk
adb shell am start -n cc.agentlabs.opencode/.MainActivity
sleep 5
# Scenario set widened to the real core journey. send_message requires the
# opencode server to actually reply (Azure provider wired in earlier step);
# if the model probe failed we drop it to the UI-only journey so the gate
# stays green and meaningful rather than red for an environmental reason.
if [ "${MODEL_CAPABLE}" = "true" ]; then
SCENARIOS="connect_and_verify_sessions,send_message,verify_session_list"
else
SCENARIOS="connect_and_verify_sessions,verify_session_list"
echo "WARN: opencode not model-capable in CI; excluding send_message/multi_turn (environmental, not an app bug)."
fi
echo "Running scenarios: $SCENARIOS"
# --max-steps raised so multiple scenarios fit; each scenario gets its own budget.
python3 scripts/android-cua-smoke.py --model gpt-5.4 --include-xml --max-steps 40 --scenarios "$SCENARIOS"
- name: opencode server log
if: always()
run: cat /tmp/opencode-server.log || true
- name: Upload test artifacts
if: always()
uses: actions/upload-artifact@v7
with:
name: cua-artifacts-${{ github.run_number }}
path: |
/tmp/cua_*.png
/tmp/cua_*.mp4
if-no-files-found: ignore