Unsafe thread concurrency with fork
AWS::S3 is not threadsafe, and by the author's account it isn't even reusable, since most methods go through a class constant. To use it in threaded code, S3 operations have to be isolated in memory — which is where forking comes in.
def s3(key, data, bucket, opts)
begin
fork_to do
AWS::S3::Base.establish_connection!(
:access_key_id => KEY,
:secret_access_key => SECRET
)
AWS::S3::S3Object.store key, data, bucket, opts
end
rescue Timeout::Error
raise SubprocessTimedOut
end
end
def fork_to(timeout = 4)
r, w, pid = nil, nil, nil
begin
# Open pipe
r, w = IO.pipe
# Start subprocess
pid = fork do
# Child
begin
r.close
val = begin
Timeout.timeout(timeout) do
# Run block
yield
end
rescue Exception => e
e
end
w.write Marshal.dump val
w.close
ensure
# YOU SHALL NOT PASS
# Skip at_exit handlers.
exit!
end
end
# Parent
w.close
Timeout.timeout(timeout) do
# Read value from pipe
begin
val = Marshal.load r.read
rescue ArgumentError => e
# Marshal data too short
# Subprocess likely exited without writing.
raise Timeout::Error
end
# Return or raise value from subprocess.
case val
when Exception
raise val
else
return val
end
end
ensure
if pid
Process.kill "TERM", pid rescue nil
Process.kill "KILL", pid rescue nil
Process.waitpid pid rescue nil
end
r.close rescue nil
w.close rescue nil
end
end
The machinery involves a fair amount of bookkeeping: the code forks and runs a given block in a subprocess, then returns the result to the parent through a pipe, with the remainder handling timeouts and process accounting. Subprocesses tend to get tied up, leaving dangling pipes or zombies behind, and the author acknowledges weak points and race conditions in the approach — but considers it suitable for production when paired with robust retry code.
With this approach, roughly 8 S3 uploads can be kept running concurrently on a fairly busy 6-core HT Nehalem, yielding about sixfold the throughput of locking S3 operations with a mutex.



