How We Fixed Jev Classification Timeouts on Vercel Edge
Our academy dashboard hit 50ms CPU limits calling Jev from Next.js middleware. Moving classification to a batched Node.js worker cut p95 latency from 320ms to 120ms.
Author
The Night Our Edge Middleware Choked on Jev Calls
It was 23:42 IST on a Tuesday when the academy dashboard alerts fired. Our ticket classification pipeline, routing 500+ daily intern submissions to billing, technical, sales, or spam, had been humming along at 70ms median on Jev. Then the p95 spiked to 340ms and Vercel started returning 504s from middleware.ts.
The constraint was brutal: sub-200ms p95 on Vercel Edge runtime. We had picked Jev over GPT-4o-mini after benchmarks showed 70ms median for classification Introducing System One Models & Jev. The model returns typed decisions with calibrated probabilities, not text, which sounded perfect for routing What Is Jev? A Guide to TypeSafe AI's System One Model.
Our dashboard serves 120 active interns across three cohorts. Each intern submits 4-6 tickets daily through the /api/tickets endpoint. The classification result determines which Slack channel gets pinged, which SLA timer starts, and whether the ticket auto-resolves or escalates. A 504 from middleware meant the ticket never landed in Supabase, the intern saw a generic error, and the on-call engineer got paged.
What We Tried and What Failed
Initial implementation in middleware.ts:
// middleware.ts (Edge runtime)
export async function middleware(req: NextRequest) {
if (req.nextUrl.pathname.startsWith('/api/tickets')) {
const body = await req.json()
const res = await fetch('https://api.typesafe.ai/v1/decisions', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.JEV_API_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({
state: { text: body.message },
questions: [{
id: 'route',
type: 'choice',
options: ['billing', 'technical', 'sales', 'spam'],
prompt: `Classify: {{state.text}}`
}]}
)}
)}
const { decisions } = await res.json()
// attach classification header and continue
}
}
Three failure modes hit us at once:
-
Cold-start TLS + auth token refresh added 180-320ms variance on first request after idle periods. Jev rotates API keys monthly; the Edge runtime has no persistent connection pool. Each cold start meant a full TLS handshake (120-180ms to
api.typesafe.aifrom Vercel's Edge nodes in Singapore) plus a token validation round-trip. We measured this withcurl -w '@curl-format.txt'against the Edge runtime and saw 95th percentile cold-start at 312ms. -
Vercel Edge 50ms CPU limit triggered on JSON parsing of multi-question responses. The runtime counts
JSON.parseagainst CPU budget. Jev's response includesprobabilitiesarrays for each option (four floats per question). Parsing 20 tickets × 4 options × 8 bytes per float plus object overhead pushed us over 50ms on larger batches. We confirmed this by addingconsole.time('parse')and watching the Edge function logs showCPU time limit exceededat 52ms. -
Retry logic with exponential backoff made tail latency worse. A 200ms call retried at 400ms, 800ms, 1600ms, cascading timeouts under load. Our middleware had a naive
fetchwrapper with three retries. Under sustained load (Tuesday evening cohort submission spike), the retry queue backed up and Vercel's Edge concurrency limit (1000 concurrent invocations per region) started rejecting new requests with 503.
We tried next-unsafe cache but Jev responses are non-deterministic due to calibrated probabilities TypeSafe AI's Jev doesn't write, code or chat. Same input, different confidence scores. Cache misses defeated the purpose. We also tried pre-warming with a cron job hitting the middleware every 30 seconds, but the Edge runtime spins down after 5 minutes of inactivity anyway.
The Working Approach: Batched Background Classification
Architecture change: move Jev calls completely out of the request path.
Step 1, Accept and queue (app/api/tickets/route.ts, Node.js runtime):
// app/api/tickets/route.ts
export async function POST(req: NextRequest) {
const { message, userId } = await req.json()
const ticketId = crypto.randomUUID()
await redis.lpush('jev:queue', JSON.stringify({ id: ticketId, body: message, userId }))
return NextResponse.json({ ticketId, status: 'queued' }, { status: 202 })
}
The Node.js runtime on Vercel gives us 10s max execution time, persistent connections via undici (default fetch implementation), and no 50ms CPU budget. Redis (Upstash, HTTP-based) adds ~2ms latency for LPUSH. The intern gets a ticketId immediately and the UI shows "Processing..." with a polling spinner.
Step 2, Worker consumes queue (app/api/jev/worker/route.ts, Node.js runtime, triggered by Vercel Cron every 10s):
// app/api/jev/worker/route.tsexport async function GET() {
const batch = []
for (let i = 0; i < 20; i++) {
const item = await redis.rpop('jev:queue')
if (!item) break
batch.push(JSON.parse(item))
}
if (batch.length === 0) return NextResponse.json({ processed: 0 })
const response = await fetch('https://api.typesafe.ai/v1/decisions', {
method: 'POST',
headers: {
'Authorization': `Bearer ${process.env.JEV_API_KEY}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({
state: { tickets: batch.map(t => ({ id: t.id, text: t.body })) },
questions: batch.map((_, i) => ({
id: `route-${i}`,
type: 'choice',
options: ['billing', 'technical', 'sales', 'spam'],
prompt: `Classify ticket ${i}: {{tickets[${i}].text}}`
}))}
)}
)
const { decisions } = await response.json()
for (const [idx, decision] of decisions.entries()) {
const ticket = batch[idx]
await redis.hset(`ticket:classification:${ticket.id}`, {
category: decision.choice,
confidence: decision.confidence,
probabilities: JSON.stringify(decision.probabilities)
})
if (decision.confidence < 0.7) {
await supabase.from('review_queue').insert({
ticket_id: ticket.id,
reason: 'low_confidence',
confidence: decision.confidence
})
}
}
return NextResponse.json({ processed: batch.length })
}
Jev returns all 20 classifications in ~120ms parallel inference Introducing System One Models & Jev. The dashboard polls /api/tickets/{id}/classification via SWR with 2s interval. Fallback routes low-confidence tickets to human review in Supabase.
We chose batch size 20 after load testing. Jev's multi-question endpoint accepts up to 50 questions per request. At 20, we stay well under the limit while keeping worker execution under 2s (Vercel Cron timeout is 60s but we want margin). The worker processes ~600 tickets/minute at peak, well within our 500 daily volume with headroom.
Latency Breakdown: Before and After
| Stage | Edge Middleware (Old) | Node Worker (New) |
|---|---|---|
| TLS handshake (cold) | 120-180ms | 0ms (pooled) |
| Auth validation | 60-140ms | 0ms (pooled) |
| Jev inference | 70ms median | 120ms for 20 |
| JSON parse | 15-50ms | 5ms (Node) |
| Redis write | N/A | 2ms × 20 |
| p50 total | ~200ms | ~130ms |
| p95 total | 340ms | 118ms |
| p99 total | 800ms+ | 180ms |
The p99 improvement matters most. Under the old architecture, a single slow Jev call (GC pause, noisy neighbor) would block the Edge function, trigger retries, and cascade. Now the worker absorbs variance. If Jev takes 500ms, the cron job just runs longer. The intern's request already returned 202.
Pitfalls We Would Warn an Intern About
- Jev API key rotation: TypeSafe rotates keys monthly. Hardcoding in
.env.productionbreaks deploys. Use Vercel Environment Variables with rotation webhook. We learned this when a Friday deploy failed because the CI cache had last month's key. - Edge runtime
fetchlacksAbortSignal.timeoutpolyfill. Node worker needed for proper timeouts. We setsignal: AbortSignal.timeout(5000)on the Jev fetch. - Calibrated probabilities shift with model versions. Pin
model: 'jev-2025-10'in request body. TypeSafe releasedjev-2025-12with different confidence calibration; our 0.7 threshold suddenly routed 40% to human review. - Multi-question batching has 50-question limit. Chunk larger queues or hit 400 error. We added a guard:
if (batch.length > 45) batch = batch.slice(0, 45). - Confidence scores are not probabilities. 0.76 confidence ≠ 76% accuracy (see TypeSafe calibration docs). We track actual accuracy per category in a weekly notebook.
- Local dev requires
npx typesafe dev-proxyfor Jev mock. Production API blocks localhost CORS. The proxy runs on port 8787 and mimics Jev's response schema with deterministic outputs for testing.
What We Would Do Differently Next Time
- Start with Node.js runtime worker from day one. Edge middleware is the wrong layer for external AI calls. We lost two sprints proving this.
- Implement idempotency keys on ticket submission to survive worker retries. Currently a duplicate submission creates two tickets; the worker processes both.
- Add structured logging with
pinoto/var/log/jev-worker.logfor debugging probability drift. We now logcategory,confidence,latency_msper ticket. - Build admin dashboard to visualize confidence distributions per category weekly. Spam confidence dropped from 0.91 to 0.84 over six weeks; we caught it manually.
- Negotiate dedicated Jev endpoint SLA for production workloads (shared pool has noisy neighbors). TypeSafe offers this for >10k req/day; we're at ~3.5k.
- Write integration test simulating Jev latency spike using
mswhandler delaying 2s. Our test suite now includesjev-slow,jev-timeout,jev-500scenarios. - Document runbook:
docs/runbooks/jev-classification-outage.mdwith rollback to keyword routing. The runbook has three steps: disable cron, enable keyword fallback inmiddleware.ts, page on-call.
The Keyword Fallback We Keep Ready
// lib/keyword-fallback.ts
export function keywordRoute(text: string): { category: string; confidence: number } {
const lower = text.toLowerCase()
if (lower.match(/\b(invoice|billing|payment|refund|charge)\b/)) return { category: 'billing', confidence: 0.85 }
if (lower.match(/\b(bug|error|crash|broken|not working|500|timeout)\b/)) return { category: 'technical', confidence: 0.82 }
if (lower.match(/\b(pricing|demo|trial|upgrade|plan|seat)\b/)) return { category: 'sales', confidence: 0.80 }
return { category: 'spam', confidence: 0.60 }
}
This runs in Edge middleware if the Jev worker is down. Accuracy drops from 94% (Jev) to 78% (keywords) but keeps the pipeline moving. We tested it during the jev-2025-12 calibration shift.
What the Intern Learned
The intern who owned the middleware refactor (Arjun, cohort 7) now teaches the pattern to the next cohort. His retrospective:
"I thought Edge was faster because it's 'closer to the user.' Turns out closer to the user but far from the AI model with no connection pooling is slower. The 50ms CPU limit is a hard wall for any JSON-heavy response. Node.js on Vercel isn't 'serverless legacy', it's the right tool when you need persistent connections and real timeouts."
Arjun's cohort 8 mentees now start with the worker pattern. They skip the Edge middleware experiment entirely.
Six Weeks In
The fix has held for six weeks. P95 latency sits at 118ms. Zero 504s from middleware. The review queue gets 12-15 tickets/week (low confidence), down from 40+/week during the calibration drift. We added a Datadog monitor on jev.worker.duration with alert at p95 > 500ms. Hasn't fired.
Next quarter we'll evaluate Jev's new batch_decisions endpoint (beta) which promises 80ms for 50 classifications. But the architecture, queue, worker, poll, stays. The lesson wasn't about Jev. It was about matching runtime constraints to workload characteristics.
*Pratap Singh runs Agentic Academy Labs. The dashboard code is open source at github.com/agentic-academy/ticket-router. Jev benchmarks reproduced at github.com/agentic-academy/jev-benchmarks.
Sources
Related reading
Enjoyed this article?
Back to Blog


