NaturalCam/Tutorials/Long exposure
A long exposure on iOS, built from one-second frames
AVCaptureDevice will not hold a shutter open longer than a
second for any third-party app, on any iPhone. There is no entitlement
that lifts it. A genuine long exposure has to be built out of
seconds — and the way you combine them decides whether you get
a longer exposure or a quieter one, which are not the same thing.
Point a stock camera app at a night sky and it does something clever and opaque: Night Mode, Deep Fusion, a denoising model choosing pixels from several frames you never see individually. You cannot predict it, inspect it, or undo it. NaturalCam's one deliberate exception to single-frame capture works the other way — it is arithmetic on linear light, the same operation an observatory performs, and every number that went into it is one you chose.
The one-second ceiling
This is not a NaturalCam limitation, and it is worth being precise about
it because it shapes everything below. Apple's own camera app can open the
shutter for up to 30 seconds because it talks to hardware no third-party
app is given. Every app built on public AVFoundation API — on
every iPhone that has ever shipped — is capped at one second per exposure.
There is no capability to request, no entitlement to add. The cap is the
ceiling, and a real long exposure has to be assembled from several
exposures at that ceiling.
/// iOS will not hold a shutter open longer than a second for any /// third-party app, on any iPhone. So a genuine night exposure is not /// something a camera app can ask the hardware for — it has to be /// built out of seconds. static let maximumFrameSeconds: Double = 1
Two different questions with the same input
"Combine several frames" sounds like one operation. It is actually the answer to two different questions, and the distinction is the whole design:
| Sum | Average | |
|---|---|---|
| What it buys | A longer exposure | A quieter one |
| 8 one-second frames become | 8 seconds of light | Still 1 second of light |
| A moving subject | Trails — which for an aurora is the picture | Stays sharp |
| Random noise | Scales with the signal, unchanged | Falls as 1/√n |
Sum adds the light up: eight one-second frames become eight seconds of exposure, brighter, with moving subjects trailing the way they do on a real long exposure. Average keeps the brightness of a single frame and throws away the noise instead — random noise falls as the square root of the frame count while the signal stays put, so eight frames at ISO 12800 carry roughly the grain of ISO 1600. Averaging buys quiet, not light. Say otherwise and someone points it at the aurora and wonders why the picture is still black.
enum Combine: String, CaseIterable, Identifiable, Codable {
case sum
case average
}
struct Plan: Equatable {
let frameCount: Int
let frameSeconds: Double
let combine: Combine
/// The exposure this is equivalent to. Summing builds a genuinely
/// longer one; averaging does not — it buys quiet, not light.
var effectiveSeconds: Double {
combine == .sum ? totalSeconds : frameSeconds
}
/// How much of the noise survives, as a fraction. Averaging `n`
/// frames leaves `1/√n` of it; summing scales signal and noise
/// together and so leaves it where it was.
var noiseFraction: Double {
guard combine == .average, frameCount > 0 else { return 1 }
return 1 / Double(frameCount).squareRoot()
}
}
Frame counts on offer are powers of two — 2, 4, 8, 16, 32 — because the useful axis is doubling: each step is one more stop of light, or half the noise again. A count in between, 24 say, is a choice that costs a decision and buys a third of a stop nobody can see. The ladder stays short on purpose.
Why it has to happen in linear light
This is the detail that is easy to get wrong and impossible to notice once you have: light adds linearly, and display-encoded values do not. Average two gamma-encoded frames and the result is neither frame and not their mean either — it is a value gamma encoding never describes anything with. The frames have to be scene-linear before a single add happens.
/// Combines linear samples from several frames.
///
/// Linear is not a detail: light adds linearly and display-encoded
/// values do not, so averaging two gamma-encoded frames gives a result
/// that is neither frame nor their mean. The caller is responsible for
/// handing over scene-linear values.
static func combine(_ frames: [[Float]], using combine: Combine) -> [Float]? {
guard let first = frames.first else { return nil }
guard frames.allSatisfy({ $0.count == first.count }) else { return nil }
let scale = weight(for: combine, frameCount: frames.count)
var result = [Float](repeating: 0, count: first.count)
for frame in frames {
for index in 0..<first.count {
result[index] += frame[index] * scale
}
}
return result
}
Notice what weight(for:frameCount:) does before a single
pixel is touched: summing multiplies by 1 — a plain add — and averaging
divides by the count first, before adding, rather than adding
everything and dividing at the end.
static func weight(for combine: Combine, frameCount: Int) -> Float {
guard frameCount > 0 else { return 1 }
switch combine {
case .sum: return 1
case .average: return 1 / Float(frameCount)
}
}
That ordering is not a style choice. Divide-then-add keeps every partial result inside the domain the whole way through; add-then-divide lets the running total run past what a buffer can represent, and a clamp at the end does not recover a highlight that has already been lost to overflow.
The part that actually decides whether this ships
The obvious implementation — collect every frame, then combine them at the
end — does not survive contact with a real stack. A Bayer RAW is roughly
24 MB. Thirty-two of them is three quarters of a gigabyte of
Data, and the app is killed for memory well before the last
frame lands.
The less obvious failure is worse: even if you chain thirty-two Core Image
additions with CIAdditionCompositing, nothing is computed
until something forces it. Core Image builds a filter graph and evaluates
it lazily — so a naive chain holds all thirty-two source images alive in
that graph regardless, and the peak memory is exactly the same as
collecting them up front. The graph has to be rendered
somewhere in the middle, not just described.
final class FrameStackAccumulator {
private var buffers: [CVPixelBuffer] = []
private var accumulated: CIImage?
/// Develops one RAW flat and folds it into the total.
func add(dng: Data) throws {
let flat = try NaturalDeveloper.flatImage(fromDNG: dng)
let weight = FrameStack.weight(for: plan.combine, frameCount: plan.frameCount)
let scaled = flat.applyingFilter("CIColorMatrix", parameters: [
"inputRVector": CIVector(x: CGFloat(weight), y: 0, z: 0, w: 0),
"inputGVector": CIVector(x: 0, y: CGFloat(weight), z: 0, w: 0),
"inputBVector": CIVector(x: 0, y: 0, z: CGFloat(weight), w: 0),
])
let combined: CIImage
if let accumulated {
combined = scaled.applyingFilter("CIAdditionCompositing", parameters: [
kCIInputBackgroundImageKey: accumulated
])
} else {
combined = scaled
}
// Forces the graph to evaluate now, and frees everything upstream.
accumulated = try render(combined, extent: flat.extent)
frameCount += 1
}
}
Each frame is developed flat, scaled by the weight, composited onto the running total — and then rendered into a real pixel buffer before the next frame arrives. Rendering is what collapses the graph: it forces the work to happen now, rather than describing one more step for later, and lets everything that fed it go. Peak memory becomes one decoded frame plus two accumulation buffers, no matter whether the stack is 2 frames or 32.
Why two buffers, not one
Reading and writing the same CVPixelBuffer within a single
Core Image render is undefined, so the running total cannot land back
where it came from. The accumulator keeps two buffers and rotates between
them — each render reads the previous total from one and writes the new
one to the other:
private func render(_ image: CIImage, extent: CGRect) throws -> CIImage {
let buffer = try buffer(for: extent)
context.render(image, to: buffer, bounds: extent, colorSpace: workingColorSpace)
// Rotate, so the next add reads this one and writes the other.
buffers.append(buffers.removeFirst())
return CIImage(cvPixelBuffer: buffer, options: [.colorSpace: workingColorSpace as Any])
.cropped(to: extent)
}
And the buffers are half-float —
kCVPixelFormatType_64RGBAHalf — not the usual 8 bits per
channel. A summed stack deliberately runs past 1.0; that overrun is
what "brighter" means here. An 8-bit buffer would clip flat at the very
first frame, and every frame after it would add nothing but noise to a
highlight that was lost before the stack even started.
Capturing the stack: one RAW at a time, on purpose
The capture loop is deliberately sequential — one RAW captured, folded in and released, before the next is requested — rather than firing off several captures and collecting them:
func captureNext() {
guard accumulator.frameCount < plan.frameCount else {
finish(.success(StackResult(combined: accumulator.combinedImage()!, /* ... */)))
return
}
captureSingleRAW { result in
switch result {
case .success(let dng):
// Folded in now and released now. Waiting until the end
// would mean holding every frame at once.
try accumulator.add(dng: dng)
captureNext()
case .failure(let error):
// A stack that lost a frame is not the picture that was
// asked for. Better to say so than to quietly hand back
// a shorter exposure than the one on the badge.
finish(.failure(error))
}
}
}
Two decisions worth calling out. First: a frame that fails to capture fails the whole stack, rather than being silently skipped. A 32-frame stack that quietly became a 31-frame stack is not the exposure the badge promised, and a sum that is short one frame is a specific, wrong number of seconds — worse than an honest error.
Second, and easy to miss: the exposure is pinned for the entire stack and put back afterwards.
/// Pins the exposure where it is for the length of a stack, and
/// reports what to put back. Nil when the exposure was already fixed
/// by hand.
private func beginStackExposureHold() -> ExposureSetting? {
guard let device else { return nil }
if case .manual = activeExposure { return nil }
let previous = activeExposure
let held = ExposureSetting.holding(
liveISO: device.iso,
liveShutterSeconds: device.exposureDuration.seconds,
capabilities: capabilitiesSnapshot(for: device)
)
activeExposure = held
configureDevice { applyExposure(held, to: $0) }
return previous
}
Every frame has to be the same exposure or the arithmetic behind all of this stops meaning anything: adding a frame the meter chose to brighten to one it chose to darken averages two different photographs, not two samples of one. So the meter's live reading is captured once, at the first frame, and held fixed until the last.
The look, once — not thirty-two times
FrameStackAccumulator.combinedImage() hands back the summed
or averaged result still flat — no film look applied. The
look runs once, afterwards, on the finished total:
static func developStackedImageData(
combined: CIImage,
profile: FilmProfile = .natural,
cropFactor: CGFloat = 1,
aspect: AspectRatio = .fullFrame
) throws -> Data {
let renderer = FilmProfileRenderer.cached(for: profile, colorSpace: outputColorSpace)
return try encode(
aspectCropped(centreCropped(renderer.render(combined), by: cropFactor), to: aspect),
format: .heif
)
}
Applying grain per frame and then combining would average the grain itself into mush — thirty-two independent noise fields summing towards smoothness, which is the opposite of a film texture. Grain lands on the finished picture, once, the way it would on a single exposure.
One more thing a night exposure has to survive
A 32-frame stack is thirty-odd seconds of the sensor running flat out. A phone that is already warm will throttle partway through — and throttling does not fail cleanly, it just makes the later frames arrive slower than the earlier ones. For a summed exposure that means the light is not evenly spread across the time it claims to cover, silently.
/// Long bursts are refused outright once the phone is critical: at
/// that point iOS is shedding load wherever it can and the frames
/// would not be the exposure that was asked for.
static func allowsStack(_ state: State) -> Bool {
state < .critical
}
What the tests pin down
The arithmetic above is simple; the claims made about it are not, which is exactly why they are tested rather than trusted:
func testSummingBuildsAGenuinelyLongerExposure() {
let plan = FrameStack.Plan(frameCount: 8, frameSeconds: 1, combine: .sum)
XCTAssertEqual(plan.effectiveSeconds, 8)
XCTAssertEqual(plan.stopsGained, 3, accuracy: 1e-9)
}
func testAveragingBuysQuietRatherThanLight() {
let plan = FrameStack.Plan(frameCount: 8, frameSeconds: 1, combine: .average)
XCTAssertEqual(plan.totalSeconds, 8, "the phone still has to stay still for eight seconds")
XCTAssertEqual(plan.effectiveSeconds, 1, "but the exposure is still one second")
}
func testAveragingCutsNoiseBySquareRootOfTheCount() {
XCTAssertEqual(
FrameStack.Plan(frameCount: 16, frameSeconds: 1, combine: .average).noiseFraction,
0.25, accuracy: 1e-9
)
}
Combine in linear light, always. If you take one thing from this: never average or add display-encoded pixel values. Convert to scene-linear first, combine, then encode back. Skipping this step does not crash — it just quietly produces a wrong picture that looks almost right.
Render to collapse the graph, not to save the result. A lazy filter graph will happily describe a computation it never performs until you force it — and until you do, everything upstream stays alive. If a loop is meant to bound memory, something inside it has to actually render.
Where this lives in NaturalCam
Sources/Core/FrameStack.swift the plan: sum vs average, weights, noise math Sources/App/FrameStackAccumulator.swift the streaming accumulator, two half-float buffers Sources/App/CameraController.swift captureStack — sequential capture, exposure hold Sources/App/NaturalDeveloper.swift developStackedImageData — the look, applied once Sources/Core/StorageGuard.swift ThermalGuard — refusing a stack the phone can't finish Tests/NaturalCamCoreTests/FrameStackTests.swift
NaturalCam takes one Bayer RAW frame from one locked physical lens and develops it with Apple's processing switched off. A stack is the one place that rule bends — several frames instead of one — and it bends the same way the rest of the app does: nothing you cannot account for.