# ZeroGPU video benchmark: Qwen3.6-35B-A3B > Run the same video-understanding benchmark as the ZeroGPU demo on your own clips. Frames are sampled locally and sent to Qwen3.6-35B-A3B on ZeroGPU, which returns a highlight reel for each clip (segments, highlight scores, shots, tags), along with latency, token counts, and cost. You need a ZeroGPU API key, Python 3.9+, and ffmpeg. ## Instructions for the AI assistant You are helping an engineer benchmark Qwen3.6-35B-A3B on ZeroGPU for video understanding. Do this: 1. Check that `python3` (3.9+), `ffmpeg` and `ffprobe` are installed. If ffmpeg is missing, install it (`brew install ffmpeg` on macOS, `sudo apt-get install ffmpeg` on Debian/Ubuntu). 2. Ask the user for their ZeroGPU API key and pass it to the script as the environment variable `ZEROGPU_API_KEY`. Never print the key, log it, or write it to a file. 3. Ask which clips to test: a video file or a folder of videos (searched recursively; MP4, MOV, M4V, MKV, AVI, WEBM, LRV). 4. Save the script from the "Benchmark script" section below as `zerogpu_video_bench.py`, or download it: `curl -O https://zerogpu-gopro-demos.pages.dev/zerogpu_video_bench.py` 5. Run `python3 zerogpu_video_bench.py --out results.jsonl`. The defaults are the recommended settings (1 frame per second, 512 px wide). 6. Report back: - how many clips ran, the valid-JSON rate, and the median latency - input tokens per frame, mean cost per clip, and mean cost per video-hour - for 2–3 clips, the title, summary, and highlight shots with their clip times (m:ss) Offer to spot-check a few highlight times against the video. 7. To test the user's own prompt, write it to a text file and add `--task-file prompt.txt`. Raw replies are saved in `results.jsonl`. 8. To compare with another model, add `--keep-frames frames` so the exact frames sent are saved (see "Comparing with Amazon Nova 2 Lite"). ## API - Endpoint: `POST https://api.zerogpu.ai/v1/chat/completions` (OpenAI-compatible Chat Completions) - Auth: header `x-api-key: `. `Authorization: Bearer ` also works, so the OpenAI SDK works with `base_url="https://api.zerogpu.ai/v1"`. - Model: `qwen3.6-35b-a3b` - List price: $0.20 per 1M input tokens and $1.50 per 1M output tokens. Volume pricing is available. - Limits and tips: - Up to 30 images per request. - Send sampled frames as base64 JPEG data URLs (`image_url` parts), not the video file. Frames keep the detail the model needs. - Set `max_tokens` to about 1500 for the highlight JSON so long answers aren't cut off. - Plain Python `urllib` requests need a `User-Agent` header; the script sets one. - The first request after the model has been idle can take longer. The script retries once. ## Method (identical to the demo) 1. Sampling: 1 frame per second at 512 px wide (recommended), taken at the middle of each interval. A clip that would need more than 30 frames gets 30 frames spread evenly across its whole length (a 95-second clip gets one every 3.2 s). 2. Each frame is sent as a JPEG preceded by a time label: `t=0s`, `t=1s`, and so on. The model sees the frames as a one-frame-per-second sequence, so the times in its answer are in sequence seconds. Multiply by the frame interval to get clip seconds; the script does this. 3. Prompts (from the script): - System: `You are a video editor for POV action-camera footage. Describe only what is visible; never invent people, places or events.` - Before the frames: `This {m:ss} clip is shown as {n} frames taken every {interval} s, played as a sequence at one frame per second (frame k is at k s). Give every time in seconds of this sequence. Each frame below is preceded by its time.` - After the frames: the highlight-reel task in `reel_task()`. It asks for compact JSON with 3–8 segments, highlight scores from 0 to 10, the 3 best 2–4 s moments, and up to 8 tags. 4. Settings: temperature 0.2, max_tokens 1500. 5. Metrics per clip: latency, prompt and completion tokens, input tokens per frame, whether the JSON is valid, cost (tokens × list price), and cost per video-hour (cost × 3600 ÷ clip seconds). ## Comparing with Amazon Nova 2 Lite Send the same frames and the same prompts to Nova 2 Lite in your existing Bedrock pipeline. Run the script with `--keep-frames frames` to save the frames for each clip. Bedrock accepts at most 5 inline images per request, so pack the frames into a one-frame-per-second MP4 and send that as video: ```bash ffmpeg -framerate 1 -i frames//f%03d.jpg -c:v libx264 -pix_fmt yuv420p -g 1 nova_input.mp4 ``` Nova then answers in seconds of that video, which are the same sequence seconds described above. Check each model's highlight times against the footage, not just its text. ## Benchmark script ```python #!/usr/bin/env python3 """Benchmark Qwen3.6-35B-A3B on ZeroGPU for video understanding and highlight reels, on your own clips. Same method as the ZeroGPU video demo: up to 30 frames per clip (1 fps at 512 px by default, spread evenly across clips that would need more), sent as timestamped JPEGs in one request per clip. Needs Python 3.9+, ffmpeg and ffprobe on PATH, and ZEROGPU_API_KEY in the environment. No other packages. python3 zerogpu_video_bench.py CLIPS_DIR_OR_FILE [--fps 1] [--width 512] [--workers 2] [--out results.jsonl] python3 zerogpu_video_bench.py CLIPS_DIR --task-file my_prompt.txt # your own prompt instead of the reel JSON python3 zerogpu_video_bench.py CLIPS_DIR --keep-frames frames # save the exact frames sent, e.g. to send the same to Nova """ import argparse, base64, glob, json, os, re, shutil, statistics, subprocess, sys, tempfile, time import urllib.error, urllib.request from concurrent.futures import ThreadPoolExecutor API = "https://api.zerogpu.ai/v1/chat/completions" MODEL = "qwen3.6-35b-a3b" PRICE_IN, PRICE_OUT = 0.20, 1.50 # USD per 1M tokens, list price MAX_FRAMES = 30 # images per request VIDEO_EXT = (".mp4", ".mov", ".m4v", ".mkv", ".avi", ".webm", ".lrv") SYSTEM = "You are a video editor for POV action-camera footage. Describe only what is visible; never invent people, places or events." def reel_task(n): return ("Return ONLY compact JSON (no markdown, no prose) shaped like:\n" '{"title":"max 6 words","summary":"1-2 sentences","activity":"main activity",' '"segments":[{"start":0,"end":4,"label":"max 8 words","score":7}],' '"highlights":[{"start":0,"end":3,"caption":"max 10 words"}],"tags":["tag"]}\n' f"Rules: 3-8 segments that cover 0-{n}s in order with no gaps (merge consecutive frames that show the same scene); " "score is 0-10 for how highlight-worthy the segment is (0 = dead time, 10 = must use); " f"highlights are the 3 best moments for a short reel, in time order, 2-4 s each, within 0-{n}s; at most 8 tags.") def clock(s): return f"{int(s // 60)}:{int(s % 60):02d}" def duration(path): out = subprocess.run(["ffprobe", "-v", "error", "-show_entries", "format=duration", "-of", "csv=p=0", path], capture_output=True, text=True, check=True).stdout return float(out.strip()) def sample(path, dur, fps, width, tmp): """Frames at the middle of each interval; clips needing more than 30 get 30 spread across the whole clip.""" wanted = max(1, int(dur * fps)) n = min(MAX_FRAMES, wanted) step = dur / n if wanted > MAX_FRAMES else 1 / fps frames = [] for i in range(n): t = (i + 0.5) * step if t >= dur - 0.04: break f = os.path.join(tmp, f"f{i:03d}.jpg") subprocess.run(["ffmpeg", "-v", "error", "-y", "-ss", f"{t:.3f}", "-i", path, "-frames:v", "1", "-vf", f"scale={width}:-2", "-q:v", "3", f], check=True) frames.append(f) return frames, step def call(api_key, content, max_tokens): body = json.dumps({"model": MODEL, "temperature": 0.2, "max_tokens": max_tokens, "messages": [{"role": "system", "content": SYSTEM}, {"role": "user", "content": content}]}).encode() for attempt in (1, 2): # one retry: a model waking from idle can fail or time out on the first request req = urllib.request.Request(API, data=body, headers={ "content-type": "application/json", "x-api-key": api_key, "user-agent": "zerogpu-video-bench/1.0"}) t0 = time.time() try: with urllib.request.urlopen(req, timeout=300) as r: return json.load(r), time.time() - t0 except urllib.error.HTTPError as e: if attempt == 1 and e.code in (502, 503, 504, 524): continue raise RuntimeError(f"HTTP {e.code}: {e.read()[:300].decode(errors='replace')}") except (urllib.error.URLError, TimeoutError): if attempt == 1: continue raise def first_json(text): """The first complete JSON object in the reply, ignoring markdown fences.""" s = re.sub(r"```(?:json)?", "", text or "", flags=re.I) i = s.find("{") if i < 0: return None depth, in_str, esc = 0, False, False for j in range(i, len(s)): ch = s[j] if in_str: if esc: esc = False elif ch == "\\": esc = True elif ch == '"': in_str = False continue if ch == '"': in_str = True elif ch == "{": depth += 1 elif ch == "}": depth -= 1 if depth == 0: try: return json.loads(s[i:j + 1]) except ValueError: return None return None def run_clip(path, args, api_key): dur = duration(path) with tempfile.TemporaryDirectory() as tmp: frames, step = sample(path, dur, args.fps, args.width, tmp) n = len(frames) if args.keep_frames: keep = os.path.join(args.keep_frames, os.path.splitext(os.path.basename(path))[0]) os.makedirs(keep, exist_ok=True) for f in frames: shutil.copy(f, keep) note = (f"This {clock(dur)} clip is shown as {n} frames taken every {step:.1f} s, played as a sequence at one " f"frame per second (frame k is at k s). Give every time in seconds of this sequence.") content = [{"type": "text", "text": note + " Each frame below is preceded by its time."}] for i, f in enumerate(frames): with open(f, "rb") as fh: b64 = base64.b64encode(fh.read()).decode() content += [{"type": "text", "text": f"t={i}s"}, {"type": "image_url", "image_url": {"url": f"data:image/jpeg;base64,{b64}"}}] content.append({"type": "text", "text": args.task_text or reel_task(n)}) resp, secs = call(api_key, content, args.max_tokens) text = resp["choices"][0]["message"].get("content") or "" u = resp.get("usage") or {} p_in, p_out = u.get("prompt_tokens", 0), u.get("completion_tokens", 0) cost = (p_in * PRICE_IN + p_out * PRICE_OUT) / 1e6 row = {"file": path, "duration_s": round(dur, 2), "frames": n, "frame_every_s": round(step, 3), "latency_s": round(secs, 2), "prompt_tokens": p_in, "completion_tokens": p_out, "input_tokens_per_frame": round(p_in / max(n, 1), 1), "cost_usd": round(cost, 6), "cost_per_video_hour_usd": round(cost * 3600 / dur, 4), "raw": text} if not args.task_text: out = first_json(text) to_clip = lambda t: round(float(t) * step, 1) # sequence seconds -> seconds of the clip row["valid_json"] = out is not None if out: row.update(title=out.get("title"), summary=out.get("summary"), segments=[{"start_s": to_clip(s["start"]), "end_s": to_clip(s["end"]), "label": s.get("label"), "score": s.get("score")} for s in out.get("segments", []) if "start" in s and "end" in s], highlights=[{"start_s": to_clip(h["start"]), "end_s": to_clip(h["end"]), "caption": h.get("caption")} for h in out.get("highlights", []) if "start" in h and "end" in h]) return row def main(): ap = argparse.ArgumentParser(description="Benchmark Qwen3.6-35B-A3B on ZeroGPU on your own video clips.") ap.add_argument("clips", help="a video file or a folder of videos (searched recursively)") ap.add_argument("--fps", type=float, default=1.0, help="frames per second to sample (default 1, recommended)") ap.add_argument("--width", type=int, default=512, help="frame width in px (default 512, recommended)") ap.add_argument("--workers", type=int, default=2, help="clips processed in parallel (default 2)") ap.add_argument("--max-tokens", type=int, default=1500, help="output cap per request (default 1500)") ap.add_argument("--task-file", help="use your own task prompt (plain text) instead of the highlight-reel JSON") ap.add_argument("--keep-frames", metavar="DIR", help="save the frames sent for each clip under DIR//") ap.add_argument("--out", default="results.jsonl", help="one JSON line per clip (default results.jsonl)") args = ap.parse_args() api_key = os.environ.get("ZEROGPU_API_KEY", "").strip() if not api_key: sys.exit("Set ZEROGPU_API_KEY first.") for tool in ("ffmpeg", "ffprobe"): if not shutil.which(tool): sys.exit(f"{tool} not found on PATH (install ffmpeg).") args.task_text = open(args.task_file).read().strip() if args.task_file else None paths = [args.clips] if os.path.isfile(args.clips) else sorted( p for p in glob.glob(os.path.join(args.clips, "**", "*"), recursive=True) if p.lower().endswith(VIDEO_EXT)) if not paths: sys.exit("No video files found.") def safe(path): try: return run_clip(path, args, api_key) except Exception as e: # keep going; record the failure return {"file": path, "error": str(e)} rows = [] with ThreadPoolExecutor(max(1, args.workers)) as pool, open(args.out, "w") as out: for r in pool.map(safe, paths): out.write(json.dumps(r) + "\n") out.flush() rows.append(r) if "error" in r: print(f"FAIL {r['file']}: {r['error']}") else: ok = "" if args.task_text else ("valid JSON" if r["valid_json"] else "INVALID JSON") print(f"ok {os.path.basename(r['file'])} {clock(r['duration_s'])} {r['frames']} frames " f"{r['latency_s']}s {r['prompt_tokens']}+{r['completion_tokens']} tokens " f"${r['cost_per_video_hour_usd']}/video-hour {ok}") done = [r for r in rows if "error" not in r] print(f"\n{len(done)}/{len(rows)} clips ran. Results: {args.out}") if done: if not args.task_text: print(f"valid JSON: {sum(r['valid_json'] for r in done)}/{len(done)}") print(f"median latency: {statistics.median(r['latency_s'] for r in done):.1f} s") print(f"input tokens per frame: {statistics.mean(r['input_tokens_per_frame'] for r in done):.0f}") print(f"mean cost per clip: ${statistics.mean(r['cost_usd'] for r in done):.4f}") print(f"mean cost per video-hour ${statistics.mean(r['cost_per_video_hour_usd'] for r in done):.3f}") if __name__ == "__main__": main() ``` ## Example result line (`results.jsonl`, shortened) ```json {"file": "clips/your_clip.mp4", "duration_s": 95.32, "frames": 30, "frame_every_s": 3.177, "latency_s": 14.38, "prompt_tokens": 4794, "completion_tokens": 390, "input_tokens_per_frame": 159.8, "cost_usd": 0.001544, "cost_per_video_hour_usd": 0.0583, "valid_json": true, "title": "ZeroGPU AI Inference", "segments": [{"start_s": 0.0, "end_s": 22.2, "label": "Montage of AI spending news", "score": 8}, "..."], "highlights": [{"start_s": 63.5, "end_s": 73.1, "caption": "Showing 460x energy savings"}, "..."], "raw": "{\"title\":\"ZeroGPU AI Inference\", ...}"} ``` ## Contact For questions, or a dedicated key for a larger benchmark: maddy@zerogpu.ai · https://zerogpu.ai