A regional cable TV and internet provider runs its fiber-to-the-home network on GPON: optical line terminals (OLTs) in the network, and an optical network unit (ONU) in every subscriber’s home. The operators work with a web tool that lists the OLTs and ONUs, reads the live status of an ONU, manages ONU profiles, adds new units and shows the subscriber behind each one.
That tool had been running since 2014 on Python 2 and Flask. It worked, but every query to the hardware blocked the server, and the code had reached the end of its road. This project rebuilds it while the old one keeps working.
From Python 2 to an async stack#
- Backend: Python 3.13, FastAPI, asyncpg for the database and PySNMP for the hardware, behind a new versioned API.
- Frontend: Vue 3 with PrimeVue, vue-query and a typed client generated from the API contract.
- Side by side: the old API keeps running next to the new one, so every part can be switched over, or back, on its own.
Making the hardware fast and safe to ask#
The OLTs answer over SNMP, and they are slow and easy to overload. The old code opened a new SNMP session for every call, had no timeouts or retries, and silently cut off tables longer than one bulk read. I reviewed five half-finished SNMP implementations that had piled up in the repository, kept the one with the best architecture, fixed it, and put a chain of protections in front of the hardware:
- Cache-aside with short lifetimes, and a short “stale while error” window: if a device stops answering, the last known good data is served, clearly marked as stale.
- Singleflight deduplication: identical requests that arrive at the same time share one query to the device.
- A queue per OLT with a limit on parallel operations and priorities (writes before reads).
- A circuit breaker per OLT: after repeated failures the device gets a rest, and the API answers at once instead of hanging.
- Clear error semantics: an empty table, a timeout, an unavailable device, a full queue and an unknown unit are all different answers, and errors are never cached.
- A time budget you can do arithmetic on: timeouts, retries and deadlines add up to less than the request limit, and a test checks that they still do.
Results#
All numbers are relative, measured on the real hardware or on an SNMP simulator:
- Reading an ONU’s host table without the old “trigger” writes before it: 1.6–3.1× faster, depending on the OLT. That is now the default path.
- A read served from the warm cache: about 50× faster than a cold read from the device.
- 20 identical concurrent requests → exactly 1 query to the device; 100 requests on a warm cache → 0 queries.
- Large tables now come back complete, thanks to a proper continuation loop for bulk reads.
- The Docker build context shrank from hundreds of megabytes to kilobytes, and the production image runs as a non-root user.
Built to be trusted#
- Testing without hardware: an OLT simulator and an SNMP lab let the whole stack run locally, and the integration tests run against a simulated device. Tests that would touch live equipment refuse to start without an explicitly named target.
- Safe rollout: writes to the database and SNMP writes to the network are separate switches, off until enabled. A new release goes live only after its health check passes, and a single command rolls it back.
- Clean history: credentials and real network addresses were removed from the repository and its whole history, and git hooks now block them from coming back.