๐Ÿ”ฌ NanoInfer

Ultra-lightweight AI model inference ยท Quantized .nanomodel format ยท CPU-only

checking...

Load & Compress Model

Generate Image

Inspect .nanomodel

Compress Model

API Reference

Base URL: /api Endpoints: GET /api/health โ†’ server status GET /api/models โ†’ list loaded models POST /api/models/load โ†’ {"source":"org/model","quant":"int8","tag":"main"} POST /api/models/unload?tag= โ†’ unload a model GET /api/models/{tag}/info โ†’ model info GET /api/models/{tag}/tensors?limit=100&offset=0 โ†’ tensor list POST /api/models/{tag}/tensor โ†’ {"tensor_name":"..."} โ†’ tensor stats GET /api/models/{tag}/benchmark?n=10 โ†’ latency benchmark POST /api/generate โ†’ {"prompt":"...","width":256,"height":256,"steps":20,"seed":42} โ†’ returns PNG image POST /api/generate/base64 โ†’ same but returns {"image_base64":"..."} POST /api/compress โ†’ {"source":"...","output":"...","quant":"int8","sparsity":0} GET /api/inspect?path=... โ†’ inspect .nanomodel file Interactive docs: /api/docs Python usage: import requests r = requests.post("http://localhost:7860/api/models/load", json={"source":"org/model","quant":"int4"}) print(r.json()) r = requests.post("http://localhost:7860/api/generate", json={"prompt":"a sunset","width":256,"height":256}) with open("out.png","wb") as f: f.write(r.content)