SkyCruncher · data format explainer · October 2026

Frame Receipt Format

Some of this is implemented, some is designed, and some is conceptual. Numbers are estimates unless a source is given.

Section headings carry a marker: Built Designed Concept. Ideas carry Achievable or Reach.

A Seestar frame is 16.6 MB. Every star it measured fits in 2.2 MB.

A receipt is the record SkyCruncher writes when it processes one frame: a set of small JSON files that say what each stage did, plus columnar tables that list every star it measured. Turning a frame into a receipt is lossy compression. The pixels are not kept. The science is meant to survive: every star's position, brightness, error bar and evidence, in a form you can query across thousands of frames. Further on, a design rolls those receipts up into a standard star database, with galaxies and nebulae carried as object models beside it.

The idea in one picture
Today: the bank every star × frame × band is a row · 200 rows frames → stars predict Rolled up 25 rows + 10 running totals + 20 frame records frame totals frames → satellite trail a variable star: every frame kept stored row predicted, not stored running totals frame parameters
Each level stores what the level above it cannot predict. Most frame-to-frame change is shared by every star on a frame, so it becomes a few numbers per frame. A constant star becomes a few running totals. Only the surprises keep full rows: a variable star, a satellite trail, a stray event. It is the idea behind video compression: a keyframe, then only the changes.
Built7.5 : 1A Seestar S30 Pro sub (16,594,560 bytes) against its bank deposit (2,205,554 bytes, 3,760 stars). That is the receipt today.
Built~590 BCost of one star on one Bayer frame in the bank: four band rows of about 147 bytes each.
Concept140–220×How much smaller than raw FITS the two rollup designs come out after ten years, galaxy and nebula light included, at the assumptions in section 13.

Four ideas that make it work

The tour

  1. The receiptWhat the engine writes for one frame, every column of a stored star, what it costs, and five ways to read it.Built
  2. LayersA real frame as a forward model: stars and galaxy through the optics and air, plus sky glow and noise; and the same frame redrawn from its receipt alone.Concept
  3. The rollupFour levels that each store only what the level above cannot predict, and the noise model that makes it safe.DesignedConcept
  4. Extended lightGalaxies and nebulae as object models projected into every frame, and what the M 31 work showed.Concept
  5. StorageTen years of a growing community's data, four ways, on assumptions you can change.
  6. Identity & integrityOne frame's fingerprints, from the observer's disk to its receipt and the bank, and how the rollup could cite them.BuiltDesignedConcept
  7. IdeasWhere the format could go next: ideas we can build, and longer reaches.Concept
Words used on this pagethe receipt, the bank, the rollup
The receipt and the bank
receipt
A data product that records a run. Here it means the JSON records a run writes, plus the star tables they point to.
frame / sub
One exposure from the camera, usually a FITS file. A stack is many subs combined.
Arrow
Apache Arrow, an open standard for tables stored column by column, so a reader can pull one column without touching the rest.
IPC file
Arrow's on-disk format ("inter-process communication"). Files end in .arrow and open in any Arrow library without conversion.
record batch
One chunk of rows inside an Arrow file. A file can hold many, and a reader can take them one at a time.
schema metadata
Key and value text stored in the file's header. SkyCruncher puts one sorted JSON object there: the table's name, version, and the word lists that decode its codes.
normalized
Each fact is stored once, at the level it belongs to (per run, per frame, per star), and other tables point to it by key instead of copying it.
band
One color plane of the sensor. A Seestar's Bayer sensor gives four: R, G1, G2 and B. A standard band is one of an agreed photometric system, so values from different cameras compare.
digest
A SHA-256 fingerprint. The decision digest fingerprints what a stage decided; the host digest fingerprints which build decided it.
bank
SkyCruncher's long-term archive of star measurements, one row per star, per frame, per band.
The rollup
session
One night, one rig, one target.
residual
What was observed minus what the model predicted.
nuisance parameters
Numbers that describe a frame's conditions rather than any star: zero point, sky background, seeing, airmass, drift. Also called frame parameters.
session profile
The model of one session: its frame parameters, plus one summary per star.
hardware profile
What one camera unit does on every night: gain, read noise, dark current, flat field, color response, star shape, nonlinearity, hot pixels.
star database
One row per star per standard band: position, motion, brightness, color, and how it varies.
sufficient statistics
A few running sums that rebuild a mean and a scatter exactly: the count, the weighted sum of residuals, and the weighted sum of squared residuals.
exception
A measurement the model did not predict within its noise. It is stored in full, as a bank row.
audit sample
Full rows kept forever for a fixed share of frames, about 1 in 50, whatever the model says.
extended light
Light spread over an area instead of a point: galaxies, nebulae, the glow of a cluster core.
object model
A map or fitted profile of one galaxy or nebula in one band, kept once and used for every frame that sees it.
forward projection
Turning a model of the sky into the frame a given camera would record: placed with the frame's sky-to-pixel geometry, blurred by its star shape (PSF), scaled by its zero point, with its sky added.
Part one · built today
From a frame to a receipt

What the engine writes today, how stars get stored, what that costs, and how to read it without SkyCruncher's code.

01From a FITS frame to a receipt Built

The engine runs a frame through a fixed chain of stages. Each stage writes into its own folder of the run directory, <run>/<stage>/<product>. Early stages produce big working files: the decoded pixels, the color planes, every blob the detector found. Later stages shrink the frame down to the stars that were identified against the Gaia catalog and measured. The original file is never copied. The engine remembers it by path and fingerprint.

WORKING FILES · LARGE · REGENERABLE FROM THE ORIGINAL FITS sub 16.6 MB path + fingerprint PRE-PROCESS frame.arrow pixels + header BANDS bands.arrow R · G1 · G2 · B DETECT detections.arrow every blob SOLVE pose.json where it points WALK confirmed.arrow stars matched to Gaia identified stars only THE RECEIPT · SMALL · KEPT PHOTOMETRY photometry.arrow 53 columns per star·band BANK DEPOSIT frames/000000.arrow 26 columns · ~147 B/row CURVE SEGMENT curves/000000.arrow star-major copy, derived deposit seal / 256 one star's history: one seek per segment run_receipt.json one per run · run_id recipe · build · frame fingerprint · every stage's receipt, tables, schema versions, sha256, bytes, wall time, digests indexes
The top row is the working material: large, and regenerable from the original frame with the same build. The teal row is the compact record that carries the stars. The run receipt sits under everything and names each product with its SHA-256 and byte count. The curve segment is hatched because it is a derived second copy of the bank's columns, sorted by star.

02The format itself Built

A receipt has three kinds of parts. Each has one job.

PartFormatWhat it holdsExamples
Stage receiptsJSONWhat a stage decided and why: the rules it applied, the constants, refusals in words, its timing, its two digests.walk_receipt.json
photometry_receipt.json
detect_receipt.json
pose.json
The run receiptJSONAn index, not a measurement. One per run: the recipe that names the run, the build, the frame's fingerprint, and every stage in order with its products.run_receipt.json
TablesArrow IPCThe rows: stars, detections, lens terms, per-band surfaces. Each file carries its table name and version in its schema metadata.photometry.arrow
confirmed.arrow
frames/000000.arrow

The schema registry, SCHEMAS.toml, lists every table and JSON product the engine writes and the version a reader must accept. Today it lists 32 Arrow tables and 19 JSON products. A test fails when a table's version in code disagrees with the registry.

The run receipt, by field

These are the real top-level fields of skycruncher.run_receipt version 1.4.0.

// run_receipt.json (shape; values abbreviated)
{
  "product": "skycruncher.run_receipt",
  "version": "1.4.0",
  "run_id":   { "id", "recipe", "ticket_kind", "build_id", "runtime_flips", "inputs", "stages" },
  "build":    { "build_id", "commit", "toolchain", "target", "flips", "shelf_artifact_path", "engine_version" },
  "frame":    { "role", "fingerprint", "source_path_claim", "bytes", "hardware_profile_id", ... },
  "also_read": [ /* calibration frames, same shape as frame */ ],
  "rig":      { "carried", "rig_hash", "legs_present", "source", ... },
  "stages": [
    { "stage": "photometry", "order",
      "receipt_path", "receipt_schema", "receipt_version",
      "products": [ { "path", "table", "schema_version", "sha256", "bytes" } ],
      "wall_ms", "decision_digest", "host_digest", "digest_missing_reason",
      "exit_code", "timings", "peak_bytes" }
  ],
  "derived":  [ { "key", "value", "chosen_by", "rule" } ],
  "decision_digest", "decision_digest_missing_stages", "decision_digest_rule",
  "catalog_release", "run_dir", "written_at"
}

The run ID is a recipe, not a random number: sha256 over the build, the runtime switches, each input as role:fingerprint, and the stage list. Anyone holding the receipt can recompute it and ask for the same run again.

Normalized by level

Each fact lives at one level and is pointed to from the others. Part two carries this same rule one level further.

LevelWhere it livesKeyed by
Per runrun_receipt.json, one stage receipt per stagerun_id
Per frameframe_information.arrow, hardware_profile.arrow, whenwhere.arrow; in the bank, one row of tables/frames.arrow (telescope serial, shutter-open time in MJD, exposure, which estimator measured it, which terms are in its error bars)frame fingerprint, frame_ord
Per star, per frame, per bandphotometry.arrow rows; bank deposit rowsstar_id + frame_ord + band
Per starbank tables/stars.arrow; tables/matches.arrow maps a Gaia source_id and release to the starstar_id
Per nightnight_geometry.json, night_frame_fields.json, night_positions; a stack's stack_receipt.jsonnight, stack run

A bank row names its frame by a small ordinal and its star by an opaque number of SkyCruncher's own. It repeats neither the frame's facts nor the star's. Closed vocabularies such as status and reason are stored as one-byte codes, and the word list travels in the file's metadata, so a reader that meets a newer word still reads every column.

Versioning

  • Every table and JSON product carries a major.minor.patch version. Adding a column or field is a minor bump, and an older file still reads with the new column absent, never filled with a zero.
  • A change of meaning is a bump even when no column moves. A change of shape is a major bump: the bank went to 2.0.0 when its five fixed correction terms became one list column.
  • The run receipt stamps the version of every table the run wrote, and names the catalog release the star identifiers came from.
  • Version strings sit in the host digest, not the decision digest, so bumping a version does not make two identical measurements look different.

How it streams

  • Frames and color planes can be written in Arrow's stream format, many record batches, each holding a run of whole image rows. A reader holds one batch at a time. The engine switches to this path when the data to hold exceeds half of the free memory it measured at startup.
  • The bank is append-only. A deposit adds one new file per frame and never edits a row. Re-measuring a frame adds new rows stamped with the build that made them.
  • Reading one star does not load the bank. Every 256 deposits the bank seals a curve segment, rewritten star by star with eight stars per record batch. An index says which batch holds a star, so one star's history costs one seek per segment.
  • Reading one column of any table means reading that column's buffers only. That is what columnar storage buys.

03How stars get stored Built

Stars appear at two levels of detail. photometry.arrow is the full per-frame record a scientist or a debugger reads. The bank deposit is the lean archival row that light curves are built from.

The bank deposit row, every column

One row is one star, on one frame, in one band. This is the record of truth, bank schema 2.0.0.

ColumnTypeBytesMeaning
Keys
star_idUInt648The star, by SkyCruncher's own number
frame_ordUInt324The frame, by its place in the bank
bandUInt81Code into the band list in metadata
Where and how bright
x, yFloat64 ×216Measured centroid, in pixels
flux_rawFloat648Raw brightness, never corrected in place
flux_sigmaFloat648Its error bar
flux_upper_limitFloat64, nullable8For a star looked for and not found
ladderList<Float64>44The correction terms PHOTOMETRY evaluated here (five rungs in version 1)
ladder_version, ladder_appliedUInt8, Boolean1Which correction list, and whether any rung corrected this row
Evidence
significance, p_value, fdr_criticalFloat32 ×312How sure the detection is, against a false discovery rate
joint_chi_square, joint_p_value, joint_dofFloat32 ×2, UInt89Joint evidence on position, brightness and color
position_sigma, brightness_sigma, color_sigmaFloat32 ×312How far off each axis was, in sigmas
State
statusUInt81detected, guided, non-detection, not evaluated: landscape, deviant
reason, fdr_familyUInt8, nullable ×22Why a row is a non-detection; which false-discovery family judged it
completeness_boundFloat32, nullable4The measured completeness that justifies a non-detection row
saturatedBoolean⅛The star clipped the sensor
engine_buildUInt324Code for the build that measured it

Where a bank row's 146.6 bytes go

column widths from the table above · one segment per group
Keys 13 B Position and brightness 41 B Correction ladder 44 B Evidence 33 B State 11 B Header, metadata, padding ~5 B

The single widest column is the correction ladder, a list of 64-bit floats: 44 of about 147 bytes, close to a third of every row.

Adding up the column widths gives about 143 bytes a row before file framing. The real M 31 deposit measured 146.6 bytes a row, so the difference is the file header, the schema metadata and padding.

Sky position is reached through the star: matches.arrow links each star_id to a Gaia source_id and the catalog release. The walk's confirmed.arrow also carries the catalog ra_deg and dec_deg, the predicted and measured pixel positions, and the residual in arcseconds.

The full per-frame row: photometry.arrow

skycruncher.photometry version 1.5.0 has 53 columns per star per band. They fall into these groups:

GroupColumns
Identitysource_id, confirmed_index, band, status
Positionx_plane, y_plane, x_native, y_native, sigma_x, sigma_y, sigma_xy
Shapesize_t, e1, e2, angle_rad, shape_borrowed, shape_borrowed_from
Brightnesspsf_flux, psf_flux_sigma, aperture_flux, aperture_enclosed_fraction, aperture_over_psf, peak, npix_effective, snr, eddington_mag
Background and limitsbackground, background_label, flux_variance_source, full_well_adu, full_well_source
Crowdingisolation, is_anchor, anchor_reason, group_id, group_size, neighbor_correlation, neighbor_source_id, within_group_overlap_max, group_was_split, group_seam_overlap_dropped, pass_found
Not in the cataloguncataloged_axis, uncataloged_significance, uncataloged_width_z, uncataloged_nearest_catalog_px
Correctionsladder_constant, ladder_linear, ladder_radial, ladder_polynomial, ladder_thin_plate_spline, ladder_version, ladder_applied

Status words here are measured, not on plane, uncataloged candidate and uncataloged rejected. Most numbers are 64-bit floats, so this table is wide: about 460 bytes a row on the Canon test frame. It is built for diagnosis, not for archiving.

04Reduction efficiency Built

These numbers do not all use the same frame, so compare within a row, not across rows.

Bytes per star, one Seestar S30 Pro sub of M 31

3,760 stars deposited, four bands each. The hatched bar is worked out from column widths.

Frame file, shared out per star16,594,560 B ÷ 3,760
4,413 B
Bank deposit2,205,554 B ÷ 3,760
587 B
Curve segment copy83 B × 4 rows
~333 B
01,0002,0003,0004,000
WhatInputOutputRatioKept · thrown away
Bank deposit, M 31 S30 Pro 10 s sub16,594,560 B2,205,554 B
15,040 rows
7.5 : 1Keeps 3,760 stars × 4 bands. Drops every pixel. Input size read from a sub of the same camera and format.
Synthetic bank at M 66 scale, 1,817 stars × 4 bands16,594,560 B900,602 B
123.9 B/row
18.4 : 1File size measured on synthetic rows under the older 1.x schema; the ratio assumes an S30 Pro sub as input.
Curve segment for the M 31 sub16,594,560 B~1,252,080 B
83.25 B/row
~13 : 1From column widths. It is an extra derived copy kept beside the deposit, so it adds to storage rather than replacing it.
photometry.arrow, Canon test frame49,698,930 B
decoded planes
31,664,826 B
68,800 rows
1.6 : 1Keeps 53 columns for diagnosis. The ratio is against the decoded color planes, not the original raw file.
Walk receipt, M 66 stackn/a271,383 Bn/aOne stage's JSON receipt. JSON is readable but not small.
Whole run directory, M 66 stack, early run49,772,160 B403 MB0.12 : 1Everything: decoded frame, planes, 105.7 MB of detections, an 18.4 MB preview. Taken 2026-09-14 with 7 of 11 stages run. The working files are 8 times the frame, and they are not the receipt.
Arrow buffer compression (LZ4 or Zstandard)n/an/an/aNot used. Arrow IPC supports it, but the engine writes uncompressed files today. Section 14 proposes turning it on.

The ratio depends on the field. A receipt grows with the number of stars. A frame does not. On a Bayer frame each star costs about 590 bytes in the bank, so the ratio is roughly the frame's size divided by stars × 590. On an S30 Pro sub the two meet near 28,000 stars (16,594,560 ÷ 586.6). A rich Milky Way field could get there. A sparse galaxy field compresses far more than 7.5 to 1.

2.21 MBbank deposit (stars × 586.6 B)
7.5 : 1smaller than the 16.6 MB sub

The slider runs from 200 to 40,000 stars on a log scale; the M 31 sub had 3,760. The bar is the deposit as a share of the sub's bytes, and it fills near 28,000 stars.

Planned, and decided but not built

  • No picture on the receipt. Decided 2026-09-11: the receipt carries data only. If the user saves or hosts a full image, the receipt gets a link to it. The link field is not in the run receipt yet.
  • The residual map is regenerated, not stored. Decided 2026-09-11, along with landing the receipt on the user's own disk by default.
  • Upload is opt-in. Nothing leaves the user's machine until they say yes once at first run. An export carries a scrub receipt (built, version 1.0.0) that says what was removed.
  • Complete error bars in the bank. A bank row's bar carries only the flux fit's own sigma today. The brightness surface measures an excess scatter that is not in the bar yet.

05Reading a receipt Built

Nothing here needs SkyCruncher's code. The tables are standard Arrow IPC files and the receipts are plain JSON. The column names below are real; the paths follow the run directory layout.

# pip install pyarrow
import json
import pyarrow.ipc as ipc

stars = ipc.open_file("run/photometry/photometry.arrow").read_all()
print(json.loads(stars.schema.metadata[b"skycruncher"]))   # table name, version, word lists

measured = stars.filter(stars["status"] == "measured")
print(measured.select(["source_id", "band", "psf_flux", "psf_flux_sigma", "snr"]))

Engine tables keep their name, version and settings as one sorted JSON object under the metadata key skycruncher. Bank files carry the word lists that decode their one-byte codes (bands, status words, reasons, builds) in their own metadata, so a reader turns codes back into words without SkyCruncher's code.

06Not perfect, but efficient Concept

A receipt is a measurement, not a copy. A frame is what is on the sky (stars and a galaxy) pushed through the camera's optics and the night's air, plus the glow of the air itself, plus noise. Each part can be held a different way: as rows, as a few numbers per frame, as a model kept once, or not at all. Pick a term to see it alone, what it costs, and which level holds it.

Real frame of M 31. Real frame
One real 10 s sub written as a forward model, whose terms add back up to it exactly. Optics and air is not a picture of light: it is the step every star and the galaxy pass through, shown as the star shape it applies. Each "left over" number is the scatter of 8 × 8-pixel block averages after removing every term up to that one, all four color planes together, in units of each pixel's predicted noise: 1.00 is pure noise. Byte costs are per frame unless marked "once". This is a research measurement of the rollup design on one frame, not something the engine does yet.

One frame, redrawn from its receipt

The left picture is a real 10-second sub of M 31. The right picture was drawn without any of its pixels, from the receipt the Layers tab proposes: every Gaia star at its catalog position with the brightness this frame measured, drawn through this frame's optics and air, on its sky glow, with noise at the predicted level. The bottom picture is the real frame minus the drawing (without its noise), so whatever is left there is what the receipt does not hold.

The three choices add a galaxy layer. Galaxy model is M 31, M 32 and M 110 as the night's other 351 subs saw them, with the stars removed: an object model held once, pushed through this frame's optics and air, and only its brightness is new in this frame (16 bytes). Survey projection takes the DSS2 survey plates of the same sky instead, projects them through our camera, and scales them with two numbers per color plane: the version every camera could share today, but the plates are saturated in M 31's bulge. Either way the galaxy is held once, in one shared place, and is not baked into the frame.

Real frame: M 31's bright core and dusty disk run diagonally across a field of stars, with the companion galaxies M 32 to the right of the core and M 110 at the top left.Real frame
Synthetic frame, stars only: the same stars in the same places, but no galaxy.Synthetic · stars only
Real minus synthetic, stars only, in units of each pixel's own noise: the stars are gone and M 31, M 32 and M 110 remain in gold.Real minus synthetic · stars only
Noise scale
Pure noise looks the same everywhere, on bright stars and the galaxy core too.
DrawnPer frameHeld onceLeftover structureover the galaxy, against blank skyGalaxy light
Stars onlyGaia stars, this frame's brightnesses, through its optics and air52.9 kBbrightnesses, the air, 3 ring numbers per bright starGaia DR3+ 677 kB star shape, per camera15.9×0%
+ Galaxy modelM 31, M 32, M 110 from the night's other 351 subs+16 Bits brightness481 kBthe object model1.12×96%
  a smooth model insteadbulge and disk fit (the first version of this figure)+160 B200 B1.6×88%
+ Survey projectionDSS2 red and blue, through our camera+32 B8 numbers47.9 MBtwo survey plates7.0×104%
Galaxy pixels insteadevery galaxy pixel, 4 bands, 16-bit+2.29 MBnothingnone100%
1.04times each pixel's own noise is what is left within 3 pixels of the 63 bright stars in view: the drawing is at noise. Three numbers for each bright star's core (810 bytes a frame) got it there; the two green planes agree on them star by star (correlation 0.87), so they are each star's own 10 s of air.
96%of the galaxy light comes back from the object model held once plus 16 bytes in this frame, leaving 1.12 times noise. The survey plates bring back 104% with nothing fitted but a scale, but leave the saturated bulge.
2.3 MBof galaxy pixels per frame in this view if they were stored raw. The model costs 16 bytes a frame, because the galaxy is held once and shared.
The frame: M 31, ZWO Seestar S30 Pro, one 10 s sub through the IR-cut filter at gain 200, night of 2026-09-18 (shutter open 05:53 UTC on the 19th): an 800 × 600 crop of its four Bayer color planes (1080 × 1920 each, 7.3″ a pixel), shown as R, mean G and B with one asinh stretch for every picture.
How the synthetic frames were drawn
Stars: 18,660 Gaia DR3 stars brighter than G 17, at their catalog positions corrected by this frame's astrometric field, drawn through the Layers tab's optics and air: the night's star shape (9 zones per color plane, a halo out to 60 px), this frame's core correction, a shift and a core width for each star brighter than G 12, and the transparency. Stars brighter than G 13 have a brightness per color plane; fainter ones share one, with their color from Gaia. Saturated stars and close companions take their brightness from Gaia, calibrated on the night. Bright-star cores: the first drawing left 13 bright stars (G 8 to 9.6) as blue specks, each drawn 2 to 6% too bright at 2 to 3 px. A shared shape per brightness class did not predict them on held-out stars (1 to 10% better), nor did the mean of a star's nearest bright neighbors, so the shape varies star by star. Each of the 135 stars brighter than G 10.5 now carries three ring strengths (at 0, 1.5 and 3 px on its own profile), shared by the four color planes; fitted on each green plane alone, they agree at 0.87, 0.87 and 0.32. A fourth ring at 4.5 px did not agree (−0.15) and was dropped. Galaxy model: the night's other 351 subs, reprojected and stacked (17.5 times less noise than this frame), stars removed, warped to this frame's geometry and pushed through its core correction; its brightness per plane is fitted on the inner 0.7°. Survey projection: DSS2 red (for R and both greens) and blue (for B) from CDS hips2fits at 3″ a pixel, stars filled, placed with this frame's TAN-SIP solution, put on a linear scale by a monotone transfer and blurred by the star shape, then fitted per plane with a scale and an offset on the inner galaxy. Leftover structure is the scatter of 8 × 8-pixel block means over the galaxy divided by the same on blank sky, so 1.0 means noise only. Noise scale of the difference: by default the difference is shown in units of each pixel's own predicted noise, the sky's plus the photon noise of the stars and the galaxy, so pure noise looks the same everywhere. On the sky's noise scale a bright star's core shows its own photon noise, up to 13 times the sky's, as a speck, and M 31's bulge as a patch of louder speckle; within 3 pixels of the 63 bright stars in view the leftover is 2.26 times the sky's noise but 1.04 times their own. Still left: a close pair of G 13.5 stars at the top right drawn 14% too bright, and faint stars fainter than G 17, which are not drawn (drawing Gaia's stars to G 21 changed nothing measurable on this frame). In the survey option the saturated bulge and M 32 are left, and M 110 is drawn a little too bright. Synthetic noise is drawn at the predicted level of what each option draws, so the stars-only frame has no bulge noise. Saturated star cores hold no data and are drawn as noise in the differences. Pictures drawn by a research script, not the engine.
47 of 400band rows for stars brighter than G 10 carry no flux or a negative one

What the bright-star gap teaches

The first version of this figure left a dozen bright stars at full strength in the difference. The stars were in the receipt, at the right positions. Their fluxes were not: in some bands missing or negative, in others far too high, and that drawing skipped any star with an impossible flux in one band.

In those rows the joint fit placed the star in one group with an uncataloged detection sitting on top of it: the two profiles overlap by 0.99 at the median, against 0.37 for the bright rows that came out fine, so the fit cannot tell them apart and splits the light arbitrarily. Saturation was not the cause: the frame reaches its 65,535 ceiling on only 37 pixels.

For the rollup this matters most. A design that stores only what the model fails to predict would keep these broken rows as "surprises", while the brightest, most visible stars would have no usable flux. The receipt should carry bright stars on purpose: a flag, the catalog position and magnitude, and a model of the star's core and wings (with saturation where it applies), so a broken fit falls back to the catalog instead of to nothing.

What a receipt keeps

If you keep only receipts, here is what you can and cannot get back.

You can recover

  • Every identified star's position, raw brightness, error bar and correction terms, per band
  • Light curves across nights and rigs, from the bank's per-star read
  • Stars that were looked for and not found, as upper limits with the completeness that justifies them
  • The evidence behind each detection: significance, false-discovery threshold, joint chi-square
  • When, how long and on which telescope each frame was taken
  • Exactly which build, settings and inputs made each number, and a way to rerun it
  • A synthetic picture rendered from the star list, with gaps filled from the catalog (decided 2026-09-11; not in the engine yet; the figure above was drawn by a research script)
  • With Part two's extended-light layer: galaxy and nebula light, as an object model projected into each frame plus a residual map per session (section 12)

You cannot recover

  • The pixels, and so the picture itself
  • Galaxy and nebula light, in today's receipts; section 12 designs the way back
  • Structure finer than any object model holds, and per-frame pixel noise in the galaxy
  • Anything fainter than the detection limit that nobody looked for
  • A new measurement method applied after the fact, unless the original frame is still on disk
  • Satellite trails, artifacts and transients, beyond what the full table flags as uncataloged candidates
  • In the bank, per-star shape (the full photometry table keeps it)

Raw frames keep everything and cost the most. A receipt keeps the part of a frame that a star survey needs, in a layout built for asking questions across thousands of frames. The trade only works if the measurement is good, which is why every row carries its error bar, its evidence and the build that made it.

One more limit, stated in the code itself: the first community bank's error bars were 20 to 50 times smaller than the scatter the points actually showed, because only one term went into them. The bank now names which terms are in each frame's bars, so a reader can tell a small bar from an incomplete one. Part two depends on getting this right.

Part two · design
From receipts to a standard star database

A design for storing far less per frame once the models are good enough to predict most of what a frame will show.

07Four levels, each storing the surprise DesignedConcept

Each level stores what the level above it cannot predict. A frame becomes what the session model predicts, plus whatever surprised it. A session becomes what the hardware profile and the star database predict, plus whatever surprised them. Once the models are good, nearly every star on nearly every frame is unsurprising. Then you keep a few running totals per star instead of a row per measurement, and full detail only for the surprises: variables, transients, artifacts. Galaxies and nebulae follow the same rule through object models, which sit beside the star database (section 12).

LEVELWHAT IT STORESROWS AT THE DEFAULT ASSUMPTIONS · LOG SCALE Object models CONCEPT · object × band A map or fitted profile of each galaxy and nebula, per band, projected into every frame through the levels below 2,000 maps, in all Star database CONCEPT · star × standard band Position and motion (mostly from Gaia), brightness in standard bands, color, period and amplitude if it varies 8 million, in all Hardware profile CONCEPT · unit × band Gain, read noise, dark current against temperature, flat field, color response, star shape, nonlinearity, hot pixels 6,400 a year Session DESIGNED · night × rig × target A handful of frame parameters per frame, and one summary per star: mean, scatter, number of frames, goodness of fit 120 million a year Frame BUILT · today's bank Every star, every band, every frame as a row, about 590 bytes a star on a Seestar frame, the noise included 36 billion a year side by side: together they describe what is in the sky predicts surprises predicts surprises predicts surprises 11,0001 million1 billion
Bars use a log scale: each step of 1,000 is the same length. Object models sit beside the star database: one says where the galaxies and nebulae are and how they look, the other the stars. The hardware level is tiny because a camera behaves the same way every night. The frame level is huge because it stores every measurement, including the part every other level could have predicted.

1 · Frame Built

Today's bank. Every star, every band, every frame gets a row, about 590 bytes per star on a Seestar frame. Everything is kept, including the noise.

2 · Session Designed

One night, one rig, one target: the same ~3,000 stars in ~300 subs. Most frame-to-frame change is shared by every star on a frame: transparency and zero point, sky background, seeing, airmass, tracking drift and field rotation. Each frame becomes a handful of nuisance parameters. Each star becomes one summary: mean brightness, scatter, number of frames, goodness of fit. A constant star's 300 rows collapse to one. A star whose wiggle exceeds what the noise model predicts keeps its per-frame rows, because that is signal.

3 · Hardware Concept

One camera unit across all its nights. Gain, read noise, dark current against temperature, the flat field and vignetting, color response against a standard system, star shape across the field, nonlinearity and hot pixels do not change between nights. They live in the hardware profile, one per unit: an S30 Pro's telephoto and wide cameras are separate units. Once a sensor's spatial behavior is known, the session model needs fewer free numbers, so session summaries get smaller and more accurate. Contributed darks, bias frames and flats are the fastest way to build one. Section 14 lists these terms in parametric form.

4 · Star database Concept

One row per star, in a standard system: position and motion (mostly from Gaia), brightness in standard bands, color, and a description of how it varies (period, amplitude) where it does. The predicted measurement is a function of the star's values, the hardware profile, the session conditions, and where the star landed on the sensor.

5 · Object models Concept

One model per galaxy or nebula per band: a map sharper than any Seestar frame, or a fitted profile with a map for what it misses. It is projected into each frame through the same hardware and session profiles as the stars, so a frame's extended light costs a few numbers instead of pixels. Section 12 has the detail and what the M 31 work showed.

08Predict, then store the surprise Concept

For every star on every frame, the levels above make a prediction: a value and the noise it should have. If the observed value matches within that noise, nothing is stored for this frame. The star's running totals are updated instead. If it does not match, the measurement is stored in full as an exception.

Bright stars need their own path. A failed fit is not a surprise about the sky, so a bright star whose flux comes back missing, negative or far from its catalog magnitude should be predicted from the catalog and its saturation model, and the row flagged, rather than stored as an exception (section 06 shows why).

star values hardware profile session conditions where it landed object models predicted value and predicted noise observed value from this frame within the noise? yes no update running totals count · Σ w·r · Σ w·r² most measurements · ~97% store an exception row the full bank row, kept as is the surprises · ~3%
Here r is the residual (observed minus predicted) and w is its weight, one over the predicted noise squared. Object models enter because a star in front of a galaxy sits on its light, and that light changes with seeing. The shares assume a 3% exception rate, the default used in section 13.

09Running totals that add up Concept

Per star, per band, per hardware unit, the design keeps three sufficient statistics: the count, the weighted sum of residuals, and the weighted sum of squared residuals. They rebuild the mean and the scatter exactly. They also add: session totals add into hardware totals, which add into the star's global totals. The rollup is just addition, and it can be redone in any order.

Session A · unit 1 · night 1 n 300 Σr +0.90 Σr² 0.147 mean +0.003 · scatter 0.022 Session B · unit 1 · night 2 n 280 Σr −0.56 Σr² 0.126 mean −0.002 · scatter 0.021 Session C · unit 7 · night 2 n 240 Σr +1.20 Σr² 0.132 mean +0.005 · scatter 0.023 Unit 1 totals n 580 Σr +0.34 Σr² 0.273 mean +0.001 · scatter 0.022 Unit 7 totals n 240 Σr +1.20 Σr² 0.132 mean +0.005 · scatter 0.023 Star totals · band G1 n 820 Σr +1.54 Σr² 0.405 mean residual +0.002 mag scatter 0.022 mag predicted noise 0.020 mag χ² per frame 1.23 over 820 frames: a real excess add add
One star, one band, three sessions on two camera units. Weights are set to 1 here so the sums are easy to check, and r is in magnitudes. Each box's mean is Σr ÷ n and its scatter is √(Σr² ÷ n − mean²). No single frame crossed the exception line, yet the totals show 10% more scatter than the noise model predicts, which over 820 frames is a real excess. That is how low-amplitude variability is still caught.

10The tables Concept

The same normalization as today, carried one level further. Facts live at the level they belong to, and everything below points up by key.

Stars one row per star × standard band key: star_id · band ~8 M rows · grows with new stars Hardware profiles unit × band × version key: unit_id · band · version tiny · a few thousand rows Totals star × unit × band (× session) → star_id, unit_id about one row per star seen Sessions night × unit × band → unit_id tiny · ~32,000 rows a year Exceptions star × frame, where the model failed → star_id, frame_id ideally a few % of today's rows Frames nuisance parameters per frame → session_id a few dozen numbers · ~200 B Object models object × band × version key: object_id · band · version a few MB per map · grows slowly Residual maps session × object × band → session_id, object_id one stacked map per session star_id unit_id unit_id session_id star_id session_id
Arrows point from a table to the table its key refers to. Stars grow only with new stars; hardware profiles and sessions are tiny; frames hold a few dozen numbers each. Totals are about one row per star seen per unit. Exceptions are the only table that grows with every surprising measurement, and that is where the bytes go. The two gold tables carry extended light: object models are kept once per object and band, and each session adds one residual map per object it saw.

11What makes it safe Concept

The noise model is everything. "Within the predicted noise" only means something if the error bars are right. Today's bank shows the danger: the first community bank's error bars came out 20 to 50 times too small. Too small, and everything looks surprising and nothing compresses. Too large, and real variability is averaged away.

Error bars too small 38 of 48 stored · nothing compresses Error bars right 6 of 48 stored · brightening caught Error bars too large 0 of 48 stored · brightening lost
The same 48 residuals for one star. Gray is the predicted noise; filled points fall outside it and are stored as exceptions. The star brightens near the end. The three presets (the three panels when printed) are error bars too small, right and too large: only the right ones store the brightening and nothing else. The residuals are an illustration, not measurements.
  • Prove the noise model first. A hardware profile must show its noise model holds on data it was not tuned on before the system stops recording detail for that unit.
  • Keep an audit sample. Full rows for about 1 in 50 frames, forever. It catches model drift and lets a new model be tested against real rows.
  • Version the models. Exceptions and totals only mean something against the model version that produced them, so each carries it.
  • Fit hardware and stars together. The hardware is calibrated from constant stars, and star values come from calibrated hardware. Large surveys solve this by fitting both at once, and Gaia anchors most stars.
  • Model the noise on native pixels. On M 31, formal errors came out 2 to 3 times too small because debayering and resampling correlate neighboring pixels. So all modeling stays on the raw Bayer planes (section 14).
  • Admit new model terms only on held-out data. A term earns its place only if it predicts frames it was not fitted on, and outside catalogs may pull a value toward zero but never put light where none is measured. Section 14 lists both rules.
  • Never drop a bright star. Carry each one as a flag, its catalog position and magnitude, and a model of its core, wings and saturation. On the M 31 sub, 47 of 400 band rows for stars brighter than G 10 came back with no usable flux, all of them in groups where an uncataloged detection overlapped the star (section 06).
  • Know what is lost. The exact noise of each constant star on each frame. That is fine for science if the model is right, but it limits questions nobody has asked yet. That is why the audit sample and the raw-frame retention policy matter.

12Extended light: galaxies and nebulae Concept

A receipt today keeps the stars and drops the light between them. The same predict-and-store rule can bring that light back. Model each galaxy or nebula once, as an object model. Project it into every frame through the same hardware and session profiles that predict the stars. Then store only what the projection could not predict.

Observed frame what the camera recorded = Stars star database × PSF + Extended light object model, projected + Sky session sky plane + Noise predicted, not stored Residual = observed − stars − extended light − sky Per frame nothing new: the frame parameters already hold PSF, zero point and sky Per session one stacked residual map over the object, to catch model error Exceptions the circled transient, stored in full like any other surprise
A frame is the sum of what the levels predict, plus noise. Stars come from the star database, extended light from an object model, the smooth sky from the session. Subtract all three and what is left is noise and the occasional surprise. The noise is predicted, so it is not stored. The Layers tab takes a real M 31 frame apart the same way.

Tried on a real frame. The figure in section 06 does this on one M 31 sub. An object model built once from the night's other 351 subs (481 kB held once, 16 bytes per frame) brings back 96% of the galaxy light and leaves structure 1.12 times blank-sky noise. A smooth bulge-and-disk fit in its place leaves 1.6 times (1.48 with GALFIT, 1.33 with M 31's bar). The DSS2 survey plates projected through the frame (47.9 MB held once, 8 numbers per frame) take out the disk but leave the bulge, where the plates are saturated (7.0 times). Storing the same galaxy as pixels would cost 2.29 MB in every frame. One shared reference, held in one central library, plus a few bytes per frame replaces the galaxy's pixels in every frame that sees it: the object model pays for itself in the first frame, and the survey plates after about 21 frames of M 31.

How a model reaches a frame

  1. The object model. A map of the object in each band, fixed on the sky. It can be a survey image (Pan-STARRS, Hubble, IPHAS H-alpha), a layer fitted jointly from SkyCruncher's own subs, or a fitted profile (a bulge and a disk, or a Sérsic profile) with a map for what the profile misses. Bright stars inside it come from the star database, never from the map.
  2. Geometry. The frame's solved pose places the model on the sensor, pixel by pixel.
  3. Blur. The model is convolved with the frame's star shape (PSF): the hardware profile's shape across the field, widened by the session's seeing and smeared by the mount's field rotation during the sub.
  4. Band and scale. The hardware profile's color response maps the model's band onto the camera's; the frame's zero point scales it.
  5. Sky. The session's sky plane is added. The result is the predicted frame, before noise.

What is stored, and what is lost

WhereWhatWhy
Once per object and bandThe object model and its version, a few MB per mapEvery frame of that object reuses it.
Measured on M 31Object model: 481 kB once, 16 B a frame. Survey plates: 47.9 MB once, 32 B a frame. A smooth fitted model: 200 B once, 160 B a frame. Galaxy pixels: 2.29 MB a frame (one 800 × 600 view, 4 bands, 16-bit)From the section 06 figure. Pixels grow with every frame; a reference does not.
Per frameNothing newPSF, zero point, sky plane and registration offset are already frame parameters.
Per sessionOne stacked residual map over the object's footprint, per band, full or reduced resolutionModel error is the same in every frame, so stacking a night's residuals shows it while the noise averages down. A large residual map is the signal to update the model.
ExceptionsAnything that changes during the night: a supernova, a comet, a satelliteStored like any other surprise.
CostAt the section 13 defaults, about 16 GB a year plus 8 GB onceUnder a tenth of either rollup. Apart from raw FITS, it is the only way the archive keeps galaxy and nebula light.
LostPer-frame pixel noise in the galaxy, and detail finer than the model holdsA new way of measuring the galaxy needs the model and residual maps, or the raw frames.

What the M 31 V1 work showed

From the effort to recover Hubble's Cepheid V1 in Andromeda. Flux is in f19, the light of a G = 19 star.

  • Projection works at scale. A scene template (Hubble for structure under about 20″, Pan-STARRS for larger scales, Gaia and Pan-STARRS for the 10 brightest stars) was projected through each sub's own star shape, zero point and sky plane for 5,388 subs in 20 minutes. It explained 91 to 92% of the variance in each window, leaving night-mean residuals 1.07 to 1.19 times white noise.
  • What is left is noise. After subtracting the Pan-STARRS scene from the group's deep stack, 98% of the remainder was unstructured noise. In our own 16.4-hour stack the static part is 0.07 f19, and within one telescope errors fall as one over the square root of time.
  • A public survey is sharp enough. At Seestar resolution a Hubble template ties Pan-STARRS within 1 to 2%, no template beat Pan-STARRS, and the template is about 5% of the noise.
  • Projection ties difference imaging. Both sit at each night's noise floor; the forward model is at 0.96 of a model-free floor.
  • Saturated maps fail. Bright stars must come from Gaia or Pan-STARRS, not from the image.
  • The star shape matters. A Moffat profile misread point sources by a per-frame factor of 21 to 24% rms on S30 units, which blank-site tests could not see. A measured night PSF roughly halved it.
  • Geometry must be exact. Cutouts sat 0.4″ east of Gaia, and an old resampling spline added edge errors of up to 1.15 px.
  • Galaxy light under a star changes with seeing. Neighbor light in V1's aperture went from −0.7 to +0.3 f19 between a 10.5″ night and a 6.9″ night (r = −0.98 with star width). Only a projected model corrects it frame by frame.

What the galaxy and nebula experiments showed

  • One shared map can serve every sub. For M 33, a single galaxy layer fixed on the sky (14.70″ per pixel, 1,476 × 1,476) was fitted jointly across subs with a coarse per-sub sensor screen: per-sub χ² of 1.0 to 1.04, at 0.024 s per sub per iteration on the GPU. But 16 to 21% of bright-star light leaked into the layer, and a radial flat-field error could not be told from a round galaxy. Stored as 16-bit integers (relative error under 0.03%), such a map is about 4.4 MB per band.
  • Fitted profiles reach parity, not better. A disk plus Sérsic bulge with catalog geometry only matched a single Sérsic profile (0.85 to 0.96 of its χ²). Its spiral-arm term made edge-on galaxies 4.6 to 13.4 times worse, and bulge and disk could not be separated. GALFIT fits with up to 31 parameters beat our axisymmetric model by 2.7 to 3.9 times in χ². No tool tried has a dust term, and dust lanes dominate the NGC 3628 and M 65 residuals.
  • Deep survey images are not a free template for big galaxies. Legacy Surveys DR10, projected flux-conserving with 0.12 px registration, worked for stars and sky but over-subtracted galaxy halos by 10 to 12%, with 7 to 11σ errors in the cores. For M 33 it kept only 10 to 14% of our disk light 10 to 22′ out.
  • Nebulae need an outside anchor. A multi-scale non-negative basis recovered synthetic nebulae well (error 1.5%), but on a real NGC 7000 frame it absorbed the whole ~1,400 ADU sky pedestal: one frame cannot separate sky from nebula. IPHAS H-alpha at 1″ explains 82% of our smoothed starless red light, so it works as a shape with a free amplitude.
  • The noise is correlated. Measured noise was 1.78 ADU against 1.32 predicted, and 7.6 times white at 32 px. Star masks hid 41 to 85% of galaxy light, which is why the stars must be modeled, not masked.
  • The PSF needs the mount's rotation. An alt-az mount smears the star shape along a field-rotation arc that differs per sub.

In the engine, nothing is built yet. No galaxy or nebula model exists in the engine today. The session model's design already describes the prediction this section needs: transparency times the rendered scene, plus sky, with conditions as "a few numbers per moment, not a picture".

Next steps

Port the M 31 projection as the first object-model path Achievable

Start with what worked: a Pan-STARRS or Hubble template, or a sky-fixed layer fitted from the session's own subs, plus catalog stars, through each frame's PSF, zero point and sky. The research version ran at about 0.2 s per sub on 9.6′ cutouts around V1; full frames will cost more.

object modelsframe parameterseffort: medium

Store one residual map per session and object Achievable

Stack each session's residuals over the object's footprint and keep one map per band. It is the extended-light version of an exception, and the evidence for the next model version.

sessions · residual mapseffort: small, once projection exists

Use a measured PSF per night, not a Moffat profile Achievable

The M 31 experiments showed the Moffat shape misreads S30 frames. The hardware profile holds the shape across the field; the session adds the night's seeing.

hardware profilesessioneffort: medium

A shared object-model library across contributors Reach

Every session that sees M 31 improves one versioned model, and its residual map becomes evidence for the next version. First: model versioning, and a way to merge residual maps from rigs with different PSFs and bands into one correction.

Deconvolution from many sessions Reach

Fit one sharper model to all sessions at once, each through its own PSF, to resolve detail no single Seestar frame shows. First: per-frame PSFs and registration good to a fraction of a pixel, held-out checks that the new detail is real, and the compute to fit it.

Fitted galaxy profiles instead of survey images Reach

A bulge, a disk and a Sérsic profile with a map for what they miss would be smaller and would not depend on outside surveys. First: the experiments so far reach parity at best, cannot separate bulge from disk, and have no dust term, so a model with dust and a single disk component has to beat a survey template on held-out frames. Section 14 lists the same models as a reach.

Nebula models Reach

Emission nebulae as line maps with physical line ratios, dark nebulae as absorption in front of the stars. First: an absolute anchor such as IPHAS H-alpha or our own deep stack, because one frame cannot separate sky from nebula, and a basis that covers scales from about 50 px up to the whole frame.

13Long-horizon storage: raw FITS against receipts

If exceptions run to a few percent of rows, the cost of a star on a frame drops from about 590 bytes to a few bytes of totals plus the occasional exception row. That is plausibly one to two more orders of magnitude beyond today's 7.5 to 1. The real figure depends on how often the model is surprised. The chart below compares four ways to keep a growing community's data over ten years. Both rollups include the extended-light layer of section 12 unless you switch it off below.

Assumptions

Change any number; the chart, table and arithmetic update.

Held fixed:
  • Seestar S30 Pro sub: 16,594,560 B
  • 4 bands; a session sees the same stars all night
  • Bank row (and exception row): 146.6 B
  • Frame parameters: 200 B a frame
  • Session totals: 64 B per star per band
  • Star database: 200 B per star per band, 2 million stars
  • Totals: 32 B per star per unit per band, 3 units per star
  • Hardware profile: 1 MB per unit per version, 2 units per contributor, 4 versions a year
  • Sessions table: 200 B per session per band
  • Object models: 500 objects × 4 bands × 4 MB, kept once (about one 1,476 × 1,476 map at 16 bits, the size of the M 33 sky-fixed layer)
  • Extended-light parameters: 64 B per frame that has such light
  • No compression; 1 TB = 1012 bytes

Cumulative storage over ten years

Log scale: each gridline is ten times the one below. Hover or tap for values.

Strategy1 year3 years5 years10 yearsSmaller than raw

The arithmetic

Raw FITS
Today's bank
Session consolidation
Full rollup

What each can answer later

Raw FITS

CanEverything, including measurement methods nobody has written yet.

CannotNothing is lost. The cost is the bytes.

Today's bank

CanAny star-level question: positions, brightness, error bars and evidence on every frame.

CannotAnything that needs pixels: extended light, faint sources nobody looked for, a new way of measuring.

Session consolidation

CanEach star's mean, scatter and variability per session, plus every exception in full. With extended light on, galaxy and nebula light as object models plus a residual map per session.

CannotA constant star's value on one particular frame. With no audit sample, there is nothing to re-test a new model on.

Full rollup

CanStandard-band brightness and variability for every star across all rigs, every exception, galaxy and nebula light as object models plus residual maps, and whatever the audit sample allows.

CannotThe exact per-frame noise of each constant star outside the audit sample, detail finer than the object models hold, or pixels.

Identity and integrity BuiltDesignedConcept

Every frame SkyCruncher touches gets a fingerprint: a SHA-256 hash, a 64-character code computed from the bytes. Change one bit and the code changes completely, and nobody knows how to make two different inputs share one. The engine, the upload page, the upload server and our research tools all use this same hash. This tab follows one frame's fingerprints from the observer's disk to its receipt and the bank. Then it proposes how the rollup levels could cite the fingerprints of what they summarize, so any number in the star database traces back to the frames it came from.

SHA-256The only hash function in use: in the browser, on the upload server, in the engine and in our research tools. No BLAKE3, no xxHash.
2,880 BAll the pixel data the quick key reads: one FITS block. The rest of the key is nine header cards.
0 BWhat the demo below sends anywhere. It makes no network requests; your file stays in this tab.

Three fingerprints for one frame Built

A FITS file is a chain of HDUs (header and data units). Each is a text header of 80-character cards, then the data, padded with zeros to a whole number of 2,880-byte blocks. Three hashes cover three different slices of the file, and each answers a different question. A fourth, in the engine's stacker, covers the decoded pixels.

FingerprintWhat it coversSurvives the site fuzzWhat it is for
Frame IDthe engine's fingerprintEvery byte of the file: headers, data and padding.NoThe engine's key for a frame: the run receipt, the run ID, the bank's frame table, and duplicates within one batch.
data_sha256the upload's data hashEvery HDU's data bytes in file order, each cut to the length its cards declare (BITPIX, NAXISn, PCOUNT, GCOUNT). Padding and headers excluded.YesThe upload page's duplicate check, and the upload server's index across batches. Computed on both sides and compared.
Quick keyidentity without the pixelsNine identity cards exactly as written (INSTRUME, TELESCOP, DATE-OBS, EXPTIME, EXPOSURE, BITPIX, NAXIS1 to 3), a newline, and the first 2,880 data bytes.YesAsking the server "do you already have this frame?" before reading it off a slow drive. The server answers yes or no, never where or whose.
Plane hashin the engine's stackerThe decoded color planes: their names, geometry and samples, after a version tag.YesCatching a header-edited copy of a frame the stack already kept, which the frame ID cannot see.

Why more than one? Reading every byte off a Seestar over USB runs at about 16 MB/s. The quick key lets the upload page skip frames the server already holds after reading one header and one block. It is an identity check, not an integrity check: a change deep in the pixels leaves it alone. data_sha256 is the integrity check. The page computes it while the frame uploads, and the server computes it again from the stored copy, so the page's claim is never the only record.

Try it: fingerprint a frame in your browser

Pick a FITS file, or let the page draw a small synthetic one. The page computes all three fingerprints with exactly the upload page's definitions. Then change the frame and watch which fingerprints notice.

Live demo · runs on your device

Your file never leaves this browser. The page reads it (in 4 MiB pieces when it is over 64 MiB) and hashes it on your device with the browser's own SHA-256. The demo makes no network requests and saves nothing; the page's only outside request is the site's analytics beacon, which never sees your file.

The gap. The engine's frame ID covers the whole file, so the privacy fuzz gives an uploaded frame a new ID. To the engine, the observer's original and the uploaded copy are two frames, and the same frame run from both gets two run IDs. The decision that a frame is keyed by its bytes (2026-09-13) names this exact case as the time to revisit it: "the same observation carries two hashes." data_sha256 is the bridge, and the receipt does not carry it yet. That is the first recommendation below.

The chain of custody BuiltConcept

A chain of custody means each step records what it received, by fingerprint, so anyone can check that nothing was swapped or altered on the way. The top row is the path of one uploaded frame today. The bottom row is how the chain could continue up the rollup levels from Part two.

ONE UPLOADED FRAME, TODAY Raw frame on the observer's diskframe ID: whole filedata_sha256, quick key BUILT Privacy fuzz in the browsersite moved up to 5 kmdata bytes untouched BUILT Intake server re-hashes datachecks the page's hashindexes both keys BUILT Run receipt frame ID: fuzzed fileeach product's sha256decision digests BUILT Bank deposit frame row by frame IDrows stamp their buildidentity is the rows BUILT data_sha256 + quick key unchanged checked twice recommended: carry them on frame ID a new ID: the header changed UP THE ROLLUP: EACH LEVEL CITES A ROOT OVER WHAT IT SUMMARIZES Session one night, rig, targetroot over its frameslike stack inputs CONCEPT Hardware profile one camera unitroot over its sessionstoday: table's sha256 CONCEPT Star database one root per releaserows cite their rootsshort proof per star CONCEPT Object models galaxies and nebulaecite survey sha256sresearch code does today CONCEPT Merkle root one hash over a sortedlist of inputs: changeone input and it moves Raw frame on the observer's diskdata_sha256, quick key BUILT Privacy fuzz site moved up to 5 kmframe ID changes BUILT Intake server re-hashes dataand checks the page's BUILT Run receipt products' sha256decision digests BUILT Bank deposit frame row by frame IDidentity is the rows BUILT Session root over its frameslike stack inputs CONCEPT Hardware profile root over its sessionsone per camera unit CONCEPT Star database one root per releaseshort proof per star CONCEPT Object models cite survey sha256sfeed the star database CONCEPT data_sha256 + quick key recommended
Arrows follow the frame. Teal marks data_sha256 and the quick key: solid where they travel today, dashed where carrying them on is recommended. The frame ID breaks at the fuzz, because the fuzz rewrites the header. Colors follow the page's levels: blue for the frame level (receipt and bank), orange for session, purple for hardware, green for the star database, gold for object models. The upload steps are gray because they happen before the engine sees the frame.
  1. Raw frame Built

    On the observer's disk. The upload page reads the header and the first data block to make the quick key, then asks the server which frames it already holds and skips those. It computes data_sha256 when it reads a frame to send it. Frames with more than 512 MiB of data are not hashed.

  2. Privacy fuzz Built

    In the browser. The site cards move by a stored random offset of up to 5 km. Cards computed from the site (altitude, azimuth, airmass, hour angle and the like) are recomputed for the moved site or removed, and two HISTORY cards mark the file. The upload is built as a new header plus slices of the original file, so the data bytes are never read, copied or changed. None of the quick key's nine cards is a site card. So the quick key and data_sha256 come through unchanged, and the whole-file frame ID does not.

  3. Intake Built

    The upload server (a Cloudflare Worker, a small program that runs at Cloudflare's edge) stores the file, then hashes the stored copy's data units itself, streaming them through a digest. It writes a small hash record beside the batch and keeps two indexes in its key-value store: data_sha256 to the first stored copy, and quick key to data_sha256. A frame that is already stored is flagged, never refused or deleted. The batch manifest lists, file by file, whether the page's data_sha256 matched the server's.

  4. Run receipt BuiltDesigned

    The engine fingerprints the file it was handed: for an upload, the fuzzed copy. The run ID is a SHA-256 over the build, the runtime switches, each input as its role and frame ID, and the stage list, so the same request twice is one run. Every product is listed with its SHA-256 and byte count. The solve, walk, photometry and stack stages add a decision digest: a SHA-256 over what the stage decided, with every number rounded to the grain the measurement can resolve, so two machines that agree within that grain write the same digest.

    Designed but not yet built: digests for preprocess, bands, detect, noise model and grade. Until they exist, a full run's own rolled digest is empty. The receipt does not carry data_sha256 or the quick key.

  5. Bank deposit Built

    One file per frame that is never edited, a frame row keyed by the frame ID, and every row stamped with the build that measured it. A bank's identity is the rows it holds, never the file's bytes or their hash (decided 2026-09-15), because the bytes also depend on the Arrow library's version and settings. So a digest that cites a deposit should cover its rows, or reuse the photometry decision digest, which already identifies the measurement.

  6. Session Concept

    The session profile (designed in Part two) would cite a Merkle root over its frames, one leaf per frame: its frame ID, data_sha256, run ID and photometry decision digest. The stacker already keeps exactly this kind of leaf list: its table of inputs names every frame offered, kept or excluded, with its frame ID and plane hash.

  7. Hardware profile Concept

    Would cite the roots of the sessions it was fitted on. Today a run names its hardware profile by the SHA-256 of the profile table it wrote, which moves whenever the measured profile does.

  8. Star database Concept

    Each release would carry one root over the totals and sessions it was built from, and each star's row the roots of the sessions that fed its running totals. Because totals add up, anyone holding those session summaries can add them again and check the star's row.

  9. Object models Concept

    Each model would cite the SHA-256 of the survey images it was drawn from. Our research tools already record the address, size and SHA-256 of every survey image they download, including the DSS2 and Pan-STARRS cutouts behind the real-against-synthetic figure.

How a Merkle root works Concept

A Merkle root is one hash that stands for a whole list. Hash each item into a leaf, hash the leaves in pairs, then the pairs in pairs, until one hash is left. Change any item and the root changes. To prove one frame is in a session, show the frame's leaf and the sibling hash at each level on the way up: about 20 hashes for a million frames, 640 bytes. So a rollup does not have to keep its frames' rows to stay checkable. It keeps the root, and the list of leaves sits beside it.

Two habits keep it honest, and both have precedent. Tag what is hashed: the engine's plane hash starts with its own name and version. And mark leaves and inner nodes differently, as the public Certificate Transparency logs do (RFC 6962), so a leaf can never pass for a node.

The proposed recipe (a sketch for the engine)
// one leaf per frame in a session; the same shape one level up for sessions, profiles and releases
leaf(frame) = sha256( 0x00 || "skycruncher.session.leaf.1\n"
                      || frame_id || data_sha256 || run_id || photometry_decision_digest )
node(l, r)  = sha256( 0x01 || l || r )

root        : sort the leaves by frame_id, hash them in pairs left to right,
              let an odd one out move up unchanged, repeat until one is left
proof       : the sibling hash at each level, about log2(n) of them
// the row stores inputs_root, n_inputs and root_rule (this recipe in words, so the string is never an oracle)

Every hash SkyCruncher keeps

Beyond the frame's own three, these are the fingerprints already written today, and where.

HashWhat it covers, and whyStatus
In the engine
Product hashSHA-256 and byte count of every table and JSON file a stage writes, listed on the run receipt. It is what makes a row citable: the row names the bytes it was read from.Built
Decision digestWhat a stage decided, every number rounded to the grain the measurement can resolve, so two machines agree. The run rolls its stages' digests, in order, into one.Builtsolve, walk, photometry, stackDesignedthe other five stages
Host digestWhich build decided it: version strings and other facts about the machine, kept out of the decision digest on purpose.Built
Run IDSHA-256 over the build ID, the runtime switches, each input as role and frame ID, and the stage list. It doubles as the job ticket, and every ingredient is written beside it so anyone can recompute it.Built
Build IDSHA-256 over the commit, build switches, toolchain and target, plus the SHA-256 of the binary that actually ran.Built
Hardware profile IDFor now, the SHA-256 of the hardware profile table a run wrote, until the profiler mints a key of its own.Built
Rig hashA salted hash of the telescope, body and lens serials, so records shared beyond the observer's machine can name the rig without carrying its serials. The code that mints it exists, and every receipt still says "not carried".Designed
Reference releasesCatalog and index releases are published with a manifest of byte sizes and SHA-256, and verified as they land.Built
In the research tools
Dataset manifestSHA-256 of every file in a dataset, and one digest over the sorted lines "path, bytes, sha256": a one-level Merkle list. Data enters only through a hash-verified landing, and copies are verified against it again later.Built
FREEZE manifestSHA-256 of the engine binary, the index manifest and an experiment's scripts, written before the experiment looks at results, so a later reader knows exactly what ran.Built
Survey image manifestAddress, byte count and SHA-256 of every downloaded survey image.Built

Recommended for the receipt format Concept

Proposals, not decisions. Each one is small on its own.

  1. Carry data_sha256 and the quick key on the receipt. Add data_sha256, quick_key and a one-line data_sha256_rule to the run receipt's frame block, as a minor version bump. The engine can hash the data units in the same streaming pass that already makes the frame ID. This joins a receipt to its upload records, and to the observer's original file across the fuzz.
  2. Say when a header was rewritten. A header_edits field naming the fuzz marker cards the frame carries, such as the privacy rule's version, so a reader knows why the frame ID will not match the observer's own file.
  3. Put the bridge in the bank. Add nullable data_sha256, run_id and photometry_decision_digest columns to the bank's frame table, as a minor bump. Then a frame deposited twice, once from the original and once from the fuzzed copy, is caught, and every bank row leads back to its receipt.
  4. Finish the decision digests. Preprocess, bands, detect, noise model and grade still owe one, so a full run's rolled digest is empty. With them, one string identifies everything a run decided.
  5. Give every rollup row an inputs root. Session, hardware profile, totals and each star database release carry inputs_root, n_inputs and root_rule, with the list of leaves stored beside them, using the recipe above. The stack is the cheapest place to start, because its inputs table already holds the leaves.
  6. Fill the rig hash. The minter exists. Until a host hands it a salt, every receipt says "not carried", and the hardware level has no stable key of its own.
  7. Keep one hash function. SHA-256 runs natively in the browser, on the upload server, in the engine and in Python. The slow part is reading the drive, not hashing, so nothing here calls for a faster function such as BLAKE3.

The demo's hashing mirrors the upload page's own code line for line, and a test checks it against that code: on synthetic frames, a real Seestar sub, a multi-table index file, and frames passed through the real site fuzz.

14Ideas: where this could go Concept

One direction goes further than receipts: archive frames as forward-model parameters plus a residual quantized at the noise floor. Named model components compete to explain each pixel, and each one is admitted only if it earns its place. Storage becomes minimum description length: the cost of the model's parameters plus the cost of whatever the model leaves unexplained. That is the same rule as Part two, applied to pixels instead of star measurements. These are the ideas that bear on this page.

Achievable

A noise-floor codec for raw frames still on disk Achievable

Stabilize the noise (an Anscombe transform), round each pixel to about half its noise, then entropy-code the integers (Rice or Zstandard). Rounding to half a sigma adds about 1% to the noise (√(1 + 1/48) ≈ 1.01). Most of the gain needs only a noise model, not a model of the scene. One estimate is 4 to 8 MB for an 18.8-megapixel frame. At that bit rate an S30 Pro sub would be about 1.8 to 3.5 MB instead of 16.6 MB, so a 30-day raw window at the section 13 defaults would hold about 350 to 700 GB instead of 3.3 TB. This is untested on Seestar subs; benchmark against FITS fpack first.

raw-frame retention windoweffort: medium, a codec plus a benchmark

Keep audit frames as compressed pixels too Achievable

With the codec, pixels for 1 frame in 50 would cost about 84 to 168 GB a year at the defaults, close to the 106 GB the audit rows cost. Pixels let a new measurement method be re-run on the audit sample, which rows cannot.

audit sampleeffort: small, after the codec

Turn on Arrow buffer compression Achievable

Zstandard suits parameter tables, and LZ4 fast staging. The engine writes uncompressed Arrow today. One-byte codes and sorted keys should shrink well; 64-bit floats less so. Measure on the M 31 deposit before deciding.

bank, curve segmentstotalsexceptionseffort: small

Hardware-profile terms in parametric form Achievable

A radial vignetting polynomial, dark current that scales with exposure and temperature, a bias pedestal, hot pixels found by staying put across dithered subs, column offsets, nonlinearity and a saturation level. Each is a few numbers or a sparse list instead of a full-frame map, built from contributed darks, bias frames and flats.

hardware profileeffort: medium; pairs with the upload page's calibration frames

A concrete list of frame parameters Achievable

A low-order sky polynomial, transparency (zero point, optionally a coarse cloud grid), airmass extinction, the PSF across the field, tracking drift, and a per-band registration offset. This is what fills the 200 bytes a frame assumed in section 13.

frames tableeffort: medium

Give exceptions a cause, and store trails once Achievable

Tag each exception with what caused it. A satellite or aircraft trail becomes one row of endpoints and width instead of a row for every star it crosses. A cosmic ray is a single-sub hit. Known asteroids, checked against the Minor Planet Center's orbits, stop false variable-star claims.

exceptions · a small events tableeffort: medium

Seed variability from the AAVSO catalog, and admit it only against a control Achievable

Cross-match the star database with AAVSO's VSX catalog for known variables. A star counts as variable only when its scatter beats matched constant stars in the same session, which is the totals check of section 09.

star databasetotalseffort: small to medium

Two admission rules and native pixels Achievable

Outside catalogs may pull a value toward zero, never put light where none is measured. A new model term must improve predictions on held-out data. And all modeling stays on the raw Bayer planes, because demosaicing correlates neighboring pixels. Section 11 adopts all three.

noise modeleffort: rules plus tests

Reaches

Frames stored as formulas alone Reach

Keep only model parameters, about 50 to 350 KB a frame by one estimate, and drop the residual. This is high risk: it erases uncataloged transients and asymmetric structure. First: a model that explains every component of a frame, proven on held-out frames, with the residual kept somewhere for a while.

Tile-by-tile storage Reach

Split each frame into tiles: empty sky as formulas only, star fields as model plus quantized residual, nebula cores and anomalies at full resolution. First: the codec above, and a scene model good enough that empty-sky tiles rebuild from formulas within the noise.

Galaxy and nebula models in the frame model Reach

Sérsic galaxies, regularized grids, galactic cirrus scaled to Planck dust maps, emission and reflection nebulae. Section 12 covers this as object models. First: each model must beat a survey template on held-out frames.

Named sky causes Reach

Moonlight, twilight, city light domes, airglow and zodiacal light, each pinned by an ephemeris or a light-pollution map, instead of an anonymous sky polynomial. First: evidence that the polynomial is not enough for star photometry, and local copies of the outside datasets.

Per-region retention rules Reach

An observation plan could tell each zone of a frame how to be processed: keep raw pixels in a window around a transient candidate and summarize everywhere else. First: the codec and tile split above, and a way for the upload page to carry the rule.

Richer variability models Reach

Transit and eclipsing-binary templates, or a free amplitude per exposure, in the star database's variability description. First: enough nights per star, and template fits checked on known systems.

Covariance in the star database Reach

Export each star's parameter uncertainties and their correlations from the solver. First: the joint fit of hardware and stars from section 11 has to exist.

15Sources

Formats, schemas and byte counts are read from SkyCruncher's engine and its measurement reports as of October 2026; the code is not public yet. The M 31 frame in Part one is the author's own: one 10 s sub from a ZWO Seestar S30 Pro on the night of 2026-09-18, with the night's other 351 subs for the object model. The M 31 findings in section 12 also use data from the Seestar Collaboration's M 31 project, with permission.

  1. Gaia DR3. Star positions, brightnesses and colors. Gaia Collaboration, Vallenari, A., et al. 2023, A&A 674, A1. This work has made use of data from the European Space Agency (ESA) mission Gaia, processed by the Gaia Data Processing and Analysis Consortium (DPAC), funded by national institutions, in particular those participating in the Gaia Multilateral Agreement.
  2. Digitized Sky Survey (DSS2). The survey projection in section 06. The Digitized Sky Surveys were produced at the Space Telescope Science Institute under U.S. Government grant NAG W-2166, based on photographic data obtained using the Oschin Schmidt Telescope on Palomar Mountain and the UK Schmidt Telescope.
  3. hips2fits, a service provided by CDS, Strasbourg, used to cut the survey images to this frame's sky (HiPS: Fernique, P., et al. 2015, A&A 578, A114).
  4. Pan-STARRS1. Templates in section 12. Chambers, K. C., et al. 2016, arXiv:1612.05560.
  5. Galaxy fitters: GALFIT (Peng, C. Y., et al. 2002, AJ 124, 266; 2010, AJ 139, 2097), IMFIT (Erwin, P. 2015, ApJ 799, 226) and AstroPhot (Stone, C. J., et al. 2023, MNRAS 525, 6377), for the smooth-model comparison rows.
  6. The AAVSO International Variable Star Index (VSX), named in section 14. Watson, C. L., Henden, A. A., and Price, A. 2006, SASS 25, 47.
  7. Certificate Transparency (RFC 6962), for the hash-tree leaf and node prefixes in the Identity tab.
  8. Apache Arrow IPC, the file format of the receipts and the bank.