Part 3 of 5
Three Hundred Files
A Socket Server in PHP, Restarted by Cron
The community site in this series had live chat. In 2008 that meant one of three things: a Java applet, a Flash client talking to a socket server, or polling an endpoint every few seconds and pretending.
This one used sockets, and the server is 330 lines of PHP running as a long-lived process on shared hosting.
The Loop
The constructor binds and then never returns. It calls run(), and run() is a while(1):
function run(){
while(1){
$this->acceptConnections();
$this->reportUsers();
$this->readData();
$this->sendData();
$this->checkHeartbeats();
$this->updateDB();
$this->wait();
}
}
Seven phases, in order, forever. Accept anything new, publish the user list, drain every socket, flush the outbound buffer to everyone, time out the silent, write state to the database, sleep.
The sleep is usleep(10000), ten milliseconds, so the loop runs a hundred times a second. Everything is non-blocking: the listening socket is set to non-blocking mode at startup, and socket_accept returns nothing rather than waiting when there is no connection. That is what makes a single-threaded loop able to serve many clients, and it is the same shape as an event loop written today, with the polling done by hand because there was no select wrapper worth using.
The structure is right. What it lacks is socket_select, which existed and would have let the process sleep until something actually happened rather than waking a hundred times a second to find nothing. On a box shared with other customers, that difference is the whole cost of running the thing.
Heartbeats
Detecting a client that has gone away is the hard part of holding connections open, because a browser that closed does not always close its socket.
if(time() - $this->connections[$i]->lastHeartbeatSent >= 5){
$this->connections[$i]->lastHeartbeatSent = time();
$this->connections[$i]->write('HEARTBEAT');
}
if(time() - $this->connections[$i]->lastHeartbeatReceived >= 15){
$this->connections[$i]->disconnect();
}
Send every five seconds, hang up after fifteen of silence. Three missed beats before eviction, which is the standard ratio and gives one lost packet and one slow response worth of tolerance before a live user gets dropped.
This is the piece of the daemon I would keep unchanged. The numbers are sensible, the asymmetry between send interval and timeout is deliberate, and the state it needs is two timestamps per connection.
The Database Connection Is Checked by Using It
A process that runs for weeks will lose its database connection. MySQL closes idle connections, and in 2008 the PHP client did not reconnect on its own.
The daemon deals with this at the top of every iteration:
$sql = "SELECT * FROM images LIMIT 1";
$res = $GLOBALS['DB']->queryNoCache($sql);
if(count($res) > 0){
}
else{
$GLOBALS['DB'] = new database;
$GLOBALS['DB']->init(...);
}
Run a query that must return a row. If it comes back empty, the connection is gone, so build a new one. The empty if branch with the real work in the else is the tell that this was written by testing rather than by design, and it works.
It also runs a hundred times a second, forever. A health check on the hot path with no interval on it is a query per loop iteration, which on this loop is around eight and a half million queries a day to find out something that changes maybe twice a week. Checking on a timer, or better, checking only after a query fails, costs nothing and removes almost all of it.
Cron Is the Supervisor
Nothing here is a service. There is no init script, no systemd unit, and no process manager, because shared hosting in 2008 offered cron and nothing else.
So the daemon is started by a cron job that shells out:
$cmd = "php -q $dir/$file > $dir/$file.log & echo \$!";
$pid = exec($cmd);
Launch it in the background, capture the pid. Run that on a schedule and a crashed daemon comes back on the next tick, which is a restart policy in the same way that a smoke alarm is a fire suppression system.
The thing that stops it starting a second copy every time is a lockfile, and the lockfile logic is more careful than the rest:
$lockFile = "lockfiles/$scriptName.lock";
if(file_exists($lockFile)){
$fp = fopen($lockFile,"r");
$lockPid = fread($fp,1024);
fclose($fp);
if(is_numeric($lockPid)){
$data = shell_exec("ps ux | grep $lockPid");
Read the pid out of the lockfile, then ask the process table whether that pid is actually alive, and skip the grep process itself in the output. That is the correct handling for a stale lockfile left by a crash, which is the failure a naive file_exists check gets wrong. Somebody had been bitten by it.
Beside all of this sits a file called HOW TO FIX CRONS!!!!!!.txt. It contains three lines:
must be
maxforum:nobody
rwxrwx---
An ownership and a permission mask. That is the entire operational runbook, and it exists because the crons silently stopped whenever the permissions drifted, and the fix was not discoverable from any error message.
The Harness Came From a Softphone
The launcher does not run the chat server. It runs this:
$cmd = "php4 -q $version/softphone_daemon.php > $version/log.txt & echo \$!";
And the cron that starts the real daemons carries four commented-out entries above the live ones, for softphone_daemon.php in three directories called live, test and backup, and for a call_control_daemon.php.
So the daemon framework, the lockfile handling, the launcher, and the health checker were written for a VoIP softphone with a live, test and backup deployment, and the chat server was fitted into it afterwards. The chat is the surviving user of infrastructure built for something else, and the $version directory variable in the launcher is a leftover of a deployment scheme the chat never used.
That is what this directory is really evidence of. The chat server got heartbeats, a pid-aware lockfile, and a restart story on its first day, because none of those were written for it.
What Transfers
A non-blocking accept and a fixed-order loop is a real architecture. Seven named phases in a fixed sequence is easy to reason about and easy to instrument, and it is what an event loop looks like before the library exists. Reach for the platform's select if there is one, because the sleep is the part that does not scale.
Check liveness on failure, not on a timer, and never on every iteration. A health check in the hot path of a loop running at 100 Hz is a query storm that finds nothing. The cheapest correct version is to reconnect when a query fails.
Verify a lockfile against the process table. A pid file only means something if something confirms the pid is alive. This is the difference between a crash costing one restart and a crash costing every restart until someone notices.
Infrastructure written for one thing is the reason the next thing ships. The chat server is 330 lines because the surrounding 200 already existed. Keeping the harness when the original project dies is worth more than the project usually was.
Next: the two scoring systems, one measuring what a member made and one measuring how others responded, kept deliberately apart.