How We Fixed Jev Classification Timeouts on Vercel Edge cover image
Back to Blog
TechnologyPublished 1 October 2026· Updated 1 October 2026· 8 min read

How We Fixed Jev Classification Timeouts on Vercel Edge

Our academy dashboard hit 50ms CPU limits calling Jev from Next.js middleware. Moving classification to a batched Node.js worker cut p95 latency from 320ms to 120ms.

The Night Our Edge Middleware Choked on Jev Calls

It was 23:42 IST on a Tuesday when the academy dashboard alerts fired. Our ticket classification pipeline, routing 500+ daily intern submissions to billing, technical, sales, or spam, had been humming along at 70ms median on Jev. Then the p95 spiked to 340ms and Vercel started returning 504s from middleware.ts.

The constraint was brutal: sub-200ms p95 on Vercel Edge runtime. We had picked Jev over GPT-4o-mini after benchmarks showed 70ms median for classification Introducing System One Models & Jev. The model returns typed decisions with calibrated probabilities, not text, which sounded perfect for routing What Is Jev? A Guide to TypeSafe AI's System One Model.

Our dashboard serves 120 active interns across three cohorts. Each intern submits 4-6 tickets daily through the /api/tickets endpoint. The classification result determines which Slack channel gets pinged, which SLA timer starts, and whether the ticket auto-resolves or escalates. A 504 from middleware meant the ticket never landed in Supabase, the intern saw a generic error, and the on-call engineer got paged.

What We Tried and What Failed

Initial implementation in middleware.ts:

// middleware.ts (Edge runtime)
export async function middleware(req: NextRequest) {
  if (req.nextUrl.pathname.startsWith('/api/tickets')) {
    const body = await req.json()
    const res = await fetch('https://api.typesafe.ai/v1/decisions', {
      method: 'POST',
      headers: {
        'Authorization': `Bearer ${process.env.JEV_API_KEY}`,
        'Content-Type': 'application/json'
      },
      body: JSON.stringify({
        state: { text: body.message },
        questions: [{
          id: 'route',
          type: 'choice',
          options: ['billing', 'technical', 'sales', 'spam'],
          prompt: `Classify: {{state.text}}`
        }]}
      )}
    )}
    const { decisions } = await res.json()
    // attach classification header and continue
  }
}

Three failure modes hit us at once:

  1. Cold-start TLS + auth token refresh added 180-320ms variance on first request after idle periods. Jev rotates API keys monthly; the Edge runtime has no persistent connection pool. Each cold start meant a full TLS handshake (120-180ms to api.typesafe.ai from Vercel's Edge nodes in Singapore) plus a token validation round-trip. We measured this with curl -w '@curl-format.txt' against the Edge runtime and saw 95th percentile cold-start at 312ms.

  2. Vercel Edge 50ms CPU limit triggered on JSON parsing of multi-question responses. The runtime counts JSON.parse against CPU budget. Jev's response includes probabilities arrays for each option (four floats per question). Parsing 20 tickets × 4 options × 8 bytes per float plus object overhead pushed us over 50ms on larger batches. We confirmed this by adding console.time('parse') and watching the Edge function logs show CPU time limit exceeded at 52ms.

  3. Retry logic with exponential backoff made tail latency worse. A 200ms call retried at 400ms, 800ms, 1600ms, cascading timeouts under load. Our middleware had a naive fetch wrapper with three retries. Under sustained load (Tuesday evening cohort submission spike), the retry queue backed up and Vercel's Edge concurrency limit (1000 concurrent invocations per region) started rejecting new requests with 503.

We tried next-unsafe cache but Jev responses are non-deterministic due to calibrated probabilities TypeSafe AI's Jev doesn't write, code or chat. Same input, different confidence scores. Cache misses defeated the purpose. We also tried pre-warming with a cron job hitting the middleware every 30 seconds, but the Edge runtime spins down after 5 minutes of inactivity anyway.

The Working Approach: Batched Background Classification

Architecture change: move Jev calls completely out of the request path.

Step 1, Accept and queue (app/api/tickets/route.ts, Node.js runtime):

// app/api/tickets/route.ts
export async function POST(req: NextRequest) {
  const { message, userId } = await req.json()
  const ticketId = crypto.randomUUID()
  
  await redis.lpush('jev:queue', JSON.stringify({ id: ticketId, body: message, userId }))
  
  return NextResponse.json({ ticketId, status: 'queued' }, { status: 202 })
}

The Node.js runtime on Vercel gives us 10s max execution time, persistent connections via undici (default fetch implementation), and no 50ms CPU budget. Redis (Upstash, HTTP-based) adds ~2ms latency for LPUSH. The intern gets a ticketId immediately and the UI shows "Processing..." with a polling spinner.

Step 2, Worker consumes queue (app/api/jev/worker/route.ts, Node.js runtime, triggered by Vercel Cron every 10s):

// app/api/jev/worker/route.tsexport async function GET() {
  const batch = []
  for (let i = 0; i < 20; i++) {
    const item = await redis.rpop('jev:queue')
    if (!item) break
    batch.push(JSON.parse(item))
  }
  if (batch.length === 0) return NextResponse.json({ processed: 0 })

  const response = await fetch('https://api.typesafe.ai/v1/decisions', {
    method: 'POST',
    headers: {
      'Authorization': `Bearer ${process.env.JEV_API_KEY}`,
      'Content-Type': 'application/json'
    },
    body: JSON.stringify({
      state: { tickets: batch.map(t => ({ id: t.id, text: t.body })) },
      questions: batch.map((_, i) => ({
        id: `route-${i}`,
        type: 'choice',
        options: ['billing', 'technical', 'sales', 'spam'],
        prompt: `Classify ticket ${i}: {{tickets[${i}].text}}`
      }))}
    )}
  )

  const { decisions } = await response.json()
  
  for (const [idx, decision] of decisions.entries()) {
    const ticket = batch[idx]
    await redis.hset(`ticket:classification:${ticket.id}`, {
      category: decision.choice,
      confidence: decision.confidence,
      probabilities: JSON.stringify(decision.probabilities)
    })
    
    if (decision.confidence < 0.7) {
      await supabase.from('review_queue').insert({
        ticket_id: ticket.id,
        reason: 'low_confidence',
        confidence: decision.confidence
      })
    }
  }

  return NextResponse.json({ processed: batch.length })
}

Jev returns all 20 classifications in ~120ms parallel inference Introducing System One Models & Jev. The dashboard polls /api/tickets/{id}/classification via SWR with 2s interval. Fallback routes low-confidence tickets to human review in Supabase.

We chose batch size 20 after load testing. Jev's multi-question endpoint accepts up to 50 questions per request. At 20, we stay well under the limit while keeping worker execution under 2s (Vercel Cron timeout is 60s but we want margin). The worker processes ~600 tickets/minute at peak, well within our 500 daily volume with headroom.

Latency Breakdown: Before and After

StageEdge Middleware (Old)Node Worker (New)
TLS handshake (cold)120-180ms0ms (pooled)
Auth validation60-140ms0ms (pooled)
Jev inference70ms median120ms for 20
JSON parse15-50ms5ms (Node)
Redis writeN/A2ms × 20
p50 total~200ms~130ms
p95 total340ms118ms
p99 total800ms+180ms

The p99 improvement matters most. Under the old architecture, a single slow Jev call (GC pause, noisy neighbor) would block the Edge function, trigger retries, and cascade. Now the worker absorbs variance. If Jev takes 500ms, the cron job just runs longer. The intern's request already returned 202.

Pitfalls We Would Warn an Intern About

  • Jev API key rotation: TypeSafe rotates keys monthly. Hardcoding in .env.production breaks deploys. Use Vercel Environment Variables with rotation webhook. We learned this when a Friday deploy failed because the CI cache had last month's key.
  • Edge runtime fetch lacks AbortSignal.timeout polyfill. Node worker needed for proper timeouts. We set signal: AbortSignal.timeout(5000) on the Jev fetch.
  • Calibrated probabilities shift with model versions. Pin model: 'jev-2025-10' in request body. TypeSafe released jev-2025-12 with different confidence calibration; our 0.7 threshold suddenly routed 40% to human review.
  • Multi-question batching has 50-question limit. Chunk larger queues or hit 400 error. We added a guard: if (batch.length > 45) batch = batch.slice(0, 45).
  • Confidence scores are not probabilities. 0.76 confidence ≠ 76% accuracy (see TypeSafe calibration docs). We track actual accuracy per category in a weekly notebook.
  • Local dev requires npx typesafe dev-proxy for Jev mock. Production API blocks localhost CORS. The proxy runs on port 8787 and mimics Jev's response schema with deterministic outputs for testing.

What We Would Do Differently Next Time

  1. Start with Node.js runtime worker from day one. Edge middleware is the wrong layer for external AI calls. We lost two sprints proving this.
  2. Implement idempotency keys on ticket submission to survive worker retries. Currently a duplicate submission creates two tickets; the worker processes both.
  3. Add structured logging with pino to /var/log/jev-worker.log for debugging probability drift. We now log category, confidence, latency_ms per ticket.
  4. Build admin dashboard to visualize confidence distributions per category weekly. Spam confidence dropped from 0.91 to 0.84 over six weeks; we caught it manually.
  5. Negotiate dedicated Jev endpoint SLA for production workloads (shared pool has noisy neighbors). TypeSafe offers this for >10k req/day; we're at ~3.5k.
  6. Write integration test simulating Jev latency spike using msw handler delaying 2s. Our test suite now includes jev-slow, jev-timeout, jev-500 scenarios.
  7. Document runbook: docs/runbooks/jev-classification-outage.md with rollback to keyword routing. The runbook has three steps: disable cron, enable keyword fallback in middleware.ts, page on-call.

The Keyword Fallback We Keep Ready

// lib/keyword-fallback.ts
export function keywordRoute(text: string): { category: string; confidence: number } {
  const lower = text.toLowerCase()
  if (lower.match(/\b(invoice|billing|payment|refund|charge)\b/)) return { category: 'billing', confidence: 0.85 }
  if (lower.match(/\b(bug|error|crash|broken|not working|500|timeout)\b/)) return { category: 'technical', confidence: 0.82 }
  if (lower.match(/\b(pricing|demo|trial|upgrade|plan|seat)\b/)) return { category: 'sales', confidence: 0.80 }
  return { category: 'spam', confidence: 0.60 }
}

This runs in Edge middleware if the Jev worker is down. Accuracy drops from 94% (Jev) to 78% (keywords) but keeps the pipeline moving. We tested it during the jev-2025-12 calibration shift.

What the Intern Learned

The intern who owned the middleware refactor (Arjun, cohort 7) now teaches the pattern to the next cohort. His retrospective:

"I thought Edge was faster because it's 'closer to the user.' Turns out closer to the user but far from the AI model with no connection pooling is slower. The 50ms CPU limit is a hard wall for any JSON-heavy response. Node.js on Vercel isn't 'serverless legacy', it's the right tool when you need persistent connections and real timeouts."

Arjun's cohort 8 mentees now start with the worker pattern. They skip the Edge middleware experiment entirely.

Six Weeks In

The fix has held for six weeks. P95 latency sits at 118ms. Zero 504s from middleware. The review queue gets 12-15 tickets/week (low confidence), down from 40+/week during the calibration drift. We added a Datadog monitor on jev.worker.duration with alert at p95 > 500ms. Hasn't fired.

Next quarter we'll evaluate Jev's new batch_decisions endpoint (beta) which promises 80ms for 50 classifications. But the architecture, queue, worker, poll, stays. The lesson wasn't about Jev. It was about matching runtime constraints to workload characteristics.


*Pratap Singh runs Agentic Academy Labs. The dashboard code is open source at github.com/agentic-academy/ticket-router. Jev benchmarks reproduced at github.com/agentic-academy/jev-benchmarks.

Enjoyed this article?

Back to Blog