Nebius logo

xWord Agent

A tool-using LangGraph agent that solves crossword puzzles on Nebius Token Factory inference. Race the same open weights against Vercel AI Gateway, generate a fresh puzzle, or read how a real 13×13 completed.

Jared Werba · FDE Candidate · August 11, 2026 · werba@protonmail.com · Résumé (PDF)

0.0s ready
The agent's moves appear here when you press Solve.

Across

    Down

      Pick a puzzle and press Solve.

      60 seconds from Jared

      Speed comparison

      End-to-end time from the browser, with the same model on both services — shorter is faster. Press Race and each new comparison adds a row below the average. One race is one sample: expect the ratio to move between runs.

      Recorded average — Nebius 3.1× faster

      Nebius
      17.5s · DeepSeek V4 Pro
      Vercel
      53.7s · DeepSeek V4 Pro

      Average of 4 recorded races, 10 August 2026. The same DeepSeek V4 Pro weights ran on both services. Puzzles: two 3×3 solves, one 5×5, and one generated 5×5 solved blind.

      A real external puzzle

      The sample puzzles above are small, so the agent was pointed at a full newspaper-size crossword: the daily 13×13 from boatloadpuzzles.com, with 60 interlocking entries. The results below are recorded, because a 13×13 is not a wait-and-watch demo.

      DeepSeek V4 Pro on Nebius — completed the full grid

      The agent filled all 60 entries and chose to submit. The grid engine verified every crossing. It took 98 turns, 16.6 minutes, and 2.44 million tokens — about $4.40 of Nebius inference. A history window keeps the cost linear: without it, a run this long would cost roughly eight times more.

      Technology

      LangGraph + tools

      The model only acts through four tools — get_state, fill_slot, clear_slot, submit. LangGraph is the control plane: one node calls the model, one runs tools against the live grid. The run ends on submit, when the model stops calling tools, after a few nudges if it answers in prose, or at the turn limit. A separate deterministic Python grid engine checks every letter on fill; rejected crossings never update the board. The model proposes answers; the engine owns state. A history window keeps long runs affordable (linear in turns, not quadratic).

      Puzzle generation

      Press Generate and the demo builds a puzzle that did not exist before. Backtracking over a shipped wordlist fills a template so every entry is a real word and every crossing agrees. DeepSeek V4 Pro on Nebius then writes the clues, with a check that clues do not leak their answers. Both services race the new puzzle blind — each run gets the empty grid and clues only; the answer key is used afterward to score. Generation and solving are separate requests so each stays under the function timeout.

      Nebius Token Factory

      Nebius is the default inference path. It speaks the OpenAI chat API, so the same agent client changes only host and key. DeepSeek V4 Pro on Nebius solves every fixed sample puzzle end-to-end. In four recorded races against the same weights on Vercel AI Gateway, Nebius averaged 17.5s vs 53.7s wall-clock (~3×, often in a 2–4× range). Only tool-capable models are offered here — Nebius V4 Flash has no tools and cannot drive this agent.

      Two services, one agent

      Each dropdown model is pinned to one service: Nebius Token Factory or Vercel AI Gateway. Tools, LangGraph loop, grid engine, and puzzle stay the same; only the endpoint and model id change. Fair pairs use the same open weights (DeepSeek V4 Pro, GLM-5.2, or Kimi K2.7 Code). Solve one at a time, or press Race / Generate to fire both at once and chart client-measured wall time.

      The assignment, and how Jared extended it

      Build an AI agent capable of solving crossword puzzles.

      A language model never rewrites the board as free text. It only calls four tools — get_state, fill_slot, clear_slot, and submit — inside a multi-turn LangGraph loop. get_state returns the live grid, numbered slots, clues, and partial patterns. fill_slot proposes one answer; the grid engine checks length and every cell against crossings and rejects conflicts without updating the board. The model sees each rejection as a tool result and revises. The run ends on submit, when the model stops calling tools, or at the turn limit. The model is trusted for answers only; the engine owns grid state.

      A working implementation of the agent.

      This page is the working agent. Select a fixed sample, press Solve, and watch the log and the grid fill. On DeepSeek V4 Pro, the agent solves all four samples at 100% letter and word accuracy on live inference (Nebius and gateway paths). A separate section records a full 13×13 external daily: completed and crossing-consistent, without the publisher's answer key. You can also run the agent from the command line — the README shows how.

      A proposed evaluation methodology for measuring the quality of its solutions.

      We measure letter accuracy, word accuracy, and solved rate, averaged over repeated runs because model output varies. Baselines bracket the score: an empty grid scores 0% (floor), the answer key scores 100% (oracle ceiling), and a wordlist filler that ignores clues scores about 9% letters (structure only). The gap between that filler and the agent isolates clue understanding. We also record turns and token cost per solve. Fixed samples are small by design; the external 13×13 reports completion and engine consistency, not a keyed accuracy claim.

      A GitHub repository containing the source code and clear instructions for running the project.

      The code is public at github.com/jaredwerba/Nebius-XWord. The README shows setup in four commands. Offline tests run without an API key (grid, tools graph, generator, API, importer). Pushes to main deploy this page automatically. Imported external puzzle text is not committed — only the importer code ships.

      Beyond the assignment.

      Jared extended the brief in six ways. The page carries Nebius branding — logo, colors, and type. A live log streams every move, with a clock and ETA. A generator builds new puzzles and races both services blind. The agent runs on two services with the same open weights held constant, so latency is measured and not claimed (~3× Nebius vs gateway on recorded DeepSeek V4 Pro races). An importer reads a real 13×13 daily from the rendered page only (no encrypted answer blob), and DeepSeek V4 Pro on Nebius submitted a complete, crossing-consistent 60-slot grid.