Ask AI

Ask my resume anything.

There is no server here. A quantized language model is downloaded to your browser, compiled to WebGPU shaders, and run on your own graphics card. Your questions never leave this machine. You can pull the network cable after it loads and it will keep answering.

One-time download, then it's local forever.

The model weights are about 1.06 GB. They stream from a CDN once and are stored in your browser's cache, so every later visit starts instantly. It needs a browser with WebGPU and about 2.2 GB of GPU memory. On a machine that cannot manage that it falls back to a smaller model automatically. Integrated graphics are fine; phones are the awkward case, since they have to fetch and hold the same weights.

Qwen3.5 2B4-bit quantized
WebGPUyour hardware
0 bytessent to any server

Around a minute on a fast connection, longer on a slow one. Instant on every visit after that, and it stays loaded while you move around the site.

A small model can get details wrong. For anything that matters, the resume and LinkedIn are authoritative, or just email me. Built with WebLLM.