The whole thing runs on one box with two cores and eight gigabytes of memory. That constraint is where everything else in this post comes from. That box hosts the web app, the database, the bot, and the forwarder — the part that watches other people's servers and pulls messages out of them into yours. For a long time it was fine. Then it was not.
The scale is the part that is easy to miss. The public directory lists six source servers right now, and between them they reach about sixty-five million members across twenty-one thousand channels. I am not talking to sixty-five million people. But Discord does not know that. It opens a socket and sends me everything that happens in every channel I can see, and it never asks which ones I actually care about.
The trouble showed up in three shapes. CPU sat high and never came back down. Messages that should have forwarded in about a second were landing two to six minutes late. And one day forwarding simply stopped, with nothing in the logs to say why.
I want to be honest about the order, because it is the actual lesson. I did not know which of those mattered most. I read the code, formed a theory, fixed the wrong thing, and did it again. Only on the third pass did I stop reasoning about the code and start counting what the process was really doing.
The counting broke it open immediately. I logged every packet the gateway delivered over one window: 31,032 events. Of those, 284 were messages in channels somebody had actually asked to forward. Everything else — typing indicators, presence updates, reactions, messages in channels nobody was watching — was work the process performed and then threw away. Ninety-nine per cent of the load existed to find one per cent of the value.
The expensive part was never the network. It is that the library turns every one of those packets into a full object before my code sees it, and files that object in a cache that grows all day. Thirty-one thousand parses, thirty-one thousand allocations. My handler then looked at the channel id and returned immediately. The decision was cheap. Getting to the point of making it was not.
So the fix was to move the decision earlier — to the raw packet, before anything gets constructed. If the channel id sitting in the JSON is not one I forward from, the packet is dropped right there, and it costs a string comparison against a set. That single change is where most of the CPU went.
Second problem. For every message that did survive the filter, the code asked the database which channels were watching it, which targets those mapped to, and what keyword filters applied. That worked out to roughly three hundred and fifty queries per message. The routing table changes maybe a few times a day. There was no reason to read it off disk thousands of times an hour. It lives in memory now, gets rebuilt when somebody changes it in their panel, and a message costs one map lookup.
Third, and this is the one I would never have found by reading code: entire servers were being resynced over and over for no reason. A hundred and thirteen full syncs in six and a half minutes. Each sync fetches every channel and category on a server that may hold millions of members. The trigger was a channel-updated event, and the events were genuine — they were the join-to-create voice channels that get created and destroyed every time somebody hops into a call. A temporary voice channel that nobody forwards from was causing a full resync of a server with millions of members, hundreds of times an hour.
Two small guards fixed it. Ignore the channel types that are never synced in the first place. Then take a signature of the fields that actually matter — id, name, type, parent, position, visibility — and compare it before and after; if nothing in that list moved, do nothing. On top of that, a debounce with a ceiling, so a burst of real changes collapses into one sync instead of thirty.
Fourth was the catalog importer. It wrote each item with three queries — insert the item, read back the id it had just written, insert the link row — one after another, for every item, on every server. That is roughly eighteen hundred serial round trips to import a single page. It is now two batched multi-row upserts and one batched read: about six. The read-back was the embarrassing part. It was asking the database for something it had just handed the database.
The silent stop was separate, smaller, and took the longest to track down. A timeout was being swallowed by a catch block that logged nothing at all. Forwarding died, the process stayed up, and every health check said it was fine. There is no clever fix for that one. The catch logs now.
Where it landed: CPU is under thirty per cent — twenty-four and a half when I took the screenshot — and memory sits around forty-five. Same features, same servers, nothing cut to get there.
If any of this is worth carrying elsewhere, it is this. None of these were hard problems, and not one of them was solved by writing faster code. Every single one was work that did not need to happen at all: a packet parsed so it could be discarded, a table read from disk that never changed, a sync triggered by something nobody was watching, a row read back from a database that had just been given it. I found them by counting what the process actually did, not by reading the code and reasoning about it — which I had already tried twice, and got wrong twice. The instinct to go and optimise the slow function is usually the wrong one. The question that pays is which of this work should not exist.