Some of this is implemented, some is designed, and some is conceptual. Numbers are estimates unless a source is given.
Section headings carry a marker: Built Designed Concept. Ideas carry Achievable or Reach.
A Seestar frame is 16.6 MB. Every star it measured fits in 2.2 MB.
A receipt is the record SkyCruncher writes when it processes one frame: a set of small JSON files that say what each stage did, plus columnar tables that list every star it measured. Turning a frame into a receipt is lossy compression. The pixels are not kept. The science is meant to survive: every star's position, brightness, error bar and evidence, in a form you can query across thousands of frames. Further on, a design rolls those receipts up into a standard star database, with galaxies and nebulae carried as object models beside it.
.arrow and open in any Arrow library without conversion.What the engine writes today, how stars get stored, what that costs, and how to read it without SkyCruncher's code.
The engine runs a frame through a fixed chain of stages. Each stage writes into its own folder of the run directory, <run>/<stage>/<product>. Early stages produce big working files: the decoded pixels, the color planes, every blob the detector found. Later stages shrink the frame down to the stars that were identified against the Gaia catalog and measured. The original file is never copied. The engine remembers it by path and fingerprint.
A receipt has three kinds of parts. Each has one job.
| Part | Format | What it holds | Examples |
|---|---|---|---|
| Stage receipts | JSON | What a stage decided and why: the rules it applied, the constants, refusals in words, its timing, its two digests. | walk_receipt.json photometry_receipt.json detect_receipt.json pose.json |
| The run receipt | JSON | An index, not a measurement. One per run: the recipe that names the run, the build, the frame's fingerprint, and every stage in order with its products. | run_receipt.json |
| Tables | Arrow IPC | The rows: stars, detections, lens terms, per-band surfaces. Each file carries its table name and version in its schema metadata. | photometry.arrow confirmed.arrow frames/000000.arrow |
The schema registry, SCHEMAS.toml, lists every table and JSON product the engine writes and the version a reader must accept. Today it lists 32 Arrow tables and 19 JSON products. A test fails when a table's version in code disagrees with the registry.
These are the real top-level fields of skycruncher.run_receipt version 1.4.0.
// run_receipt.json (shape; values abbreviated) { "product": "skycruncher.run_receipt", "version": "1.4.0", "run_id": { "id", "recipe", "ticket_kind", "build_id", "runtime_flips", "inputs", "stages" }, "build": { "build_id", "commit", "toolchain", "target", "flips", "shelf_artifact_path", "engine_version" }, "frame": { "role", "fingerprint", "source_path_claim", "bytes", "hardware_profile_id", ... }, "also_read": [ /* calibration frames, same shape as frame */ ], "rig": { "carried", "rig_hash", "legs_present", "source", ... }, "stages": [ { "stage": "photometry", "order", "receipt_path", "receipt_schema", "receipt_version", "products": [ { "path", "table", "schema_version", "sha256", "bytes" } ], "wall_ms", "decision_digest", "host_digest", "digest_missing_reason", "exit_code", "timings", "peak_bytes" } ], "derived": [ { "key", "value", "chosen_by", "rule" } ], "decision_digest", "decision_digest_missing_stages", "decision_digest_rule", "catalog_release", "run_dir", "written_at" }
The run ID is a recipe, not a random number: sha256 over the build, the runtime switches, each input as role:fingerprint, and the stage list. Anyone holding the receipt can recompute it and ask for the same run again.
Each fact lives at one level and is pointed to from the others. Part two carries this same rule one level further.
| Level | Where it lives | Keyed by |
|---|---|---|
| Per run | run_receipt.json, one stage receipt per stage | run_id |
| Per frame | frame_information.arrow, hardware_profile.arrow, whenwhere.arrow; in the bank, one row of tables/frames.arrow (telescope serial, shutter-open time in MJD, exposure, which estimator measured it, which terms are in its error bars) | frame fingerprint, frame_ord |
| Per star, per frame, per band | photometry.arrow rows; bank deposit rows | star_id + frame_ord + band |
| Per star | bank tables/stars.arrow; tables/matches.arrow maps a Gaia source_id and release to the star | star_id |
| Per night | night_geometry.json, night_frame_fields.json, night_positions; a stack's stack_receipt.json | night, stack run |
A bank row names its frame by a small ordinal and its star by an opaque number of SkyCruncher's own. It repeats neither the frame's facts nor the star's. Closed vocabularies such as status and reason are stored as one-byte codes, and the word list travels in the file's metadata, so a reader that meets a newer word still reads every column.
major.minor.patch version. Adding a column or field is a minor bump, and an older file still reads with the new column absent, never filled with a zero.Stars appear at two levels of detail. photometry.arrow is the full per-frame record a scientist or a debugger reads. The bank deposit is the lean archival row that light curves are built from.
One row is one star, on one frame, in one band. This is the record of truth, bank schema 2.0.0.
| Column | Type | Bytes | Meaning |
|---|---|---|---|
| Keys | |||
| star_id | UInt64 | 8 | The star, by SkyCruncher's own number |
| frame_ord | UInt32 | 4 | The frame, by its place in the bank |
| band | UInt8 | 1 | Code into the band list in metadata |
| Where and how bright | |||
| x, y | Float64 ×2 | 16 | Measured centroid, in pixels |
| flux_raw | Float64 | 8 | Raw brightness, never corrected in place |
| flux_sigma | Float64 | 8 | Its error bar |
| flux_upper_limit | Float64, nullable | 8 | For a star looked for and not found |
| ladder | List<Float64> | 44 | The correction terms PHOTOMETRY evaluated here (five rungs in version 1) |
| ladder_version, ladder_applied | UInt8, Boolean | 1 | Which correction list, and whether any rung corrected this row |
| Evidence | |||
| significance, p_value, fdr_critical | Float32 ×3 | 12 | How sure the detection is, against a false discovery rate |
| joint_chi_square, joint_p_value, joint_dof | Float32 ×2, UInt8 | 9 | Joint evidence on position, brightness and color |
| position_sigma, brightness_sigma, color_sigma | Float32 ×3 | 12 | How far off each axis was, in sigmas |
| State | |||
| status | UInt8 | 1 | detected, guided, non-detection, not evaluated: landscape, deviant |
| reason, fdr_family | UInt8, nullable ×2 | 2 | Why a row is a non-detection; which false-discovery family judged it |
| completeness_bound | Float32, nullable | 4 | The measured completeness that justifies a non-detection row |
| saturated | Boolean | ⅛ | The star clipped the sensor |
| engine_build | UInt32 | 4 | Code for the build that measured it |
Adding up the column widths gives about 143 bytes a row before file framing. The real M 31 deposit measured 146.6 bytes a row, so the difference is the file header, the schema metadata and padding.
Sky position is reached through the star: matches.arrow links each star_id to a Gaia source_id and the catalog release. The walk's confirmed.arrow also carries the catalog ra_deg and dec_deg, the predicted and measured pixel positions, and the residual in arcseconds.
skycruncher.photometry version 1.5.0 has 53 columns per star per band. They fall into these groups:
| Group | Columns |
|---|---|
| Identity | source_id, confirmed_index, band, status |
| Position | x_plane, y_plane, x_native, y_native, sigma_x, sigma_y, sigma_xy |
| Shape | size_t, e1, e2, angle_rad, shape_borrowed, shape_borrowed_from |
| Brightness | psf_flux, psf_flux_sigma, aperture_flux, aperture_enclosed_fraction, aperture_over_psf, peak, npix_effective, snr, eddington_mag |
| Background and limits | background, background_label, flux_variance_source, full_well_adu, full_well_source |
| Crowding | isolation, is_anchor, anchor_reason, group_id, group_size, neighbor_correlation, neighbor_source_id, within_group_overlap_max, group_was_split, group_seam_overlap_dropped, pass_found |
| Not in the catalog | uncataloged_axis, uncataloged_significance, uncataloged_width_z, uncataloged_nearest_catalog_px |
| Corrections | ladder_constant, ladder_linear, ladder_radial, ladder_polynomial, ladder_thin_plate_spline, ladder_version, ladder_applied |
Status words here are measured, not on plane, uncataloged candidate and uncataloged rejected. Most numbers are 64-bit floats, so this table is wide: about 460 bytes a row on the Canon test frame. It is built for diagnosis, not for archiving.
These numbers do not all use the same frame, so compare within a row, not across rows.
3,760 stars deposited, four bands each. The hatched bar is worked out from column widths.
| What | Input | Output | Ratio | Kept · thrown away |
|---|---|---|---|---|
| Bank deposit, M 31 S30 Pro 10 s sub | 16,594,560 B | 2,205,554 B 15,040 rows | 7.5 : 1 | Keeps 3,760 stars × 4 bands. Drops every pixel. Input size read from a sub of the same camera and format. |
| Synthetic bank at M 66 scale, 1,817 stars × 4 bands | 16,594,560 B | 900,602 B 123.9 B/row | 18.4 : 1 | File size measured on synthetic rows under the older 1.x schema; the ratio assumes an S30 Pro sub as input. |
| Curve segment for the M 31 sub | 16,594,560 B | ~1,252,080 B 83.25 B/row | ~13 : 1 | From column widths. It is an extra derived copy kept beside the deposit, so it adds to storage rather than replacing it. |
| photometry.arrow, Canon test frame | 49,698,930 B decoded planes | 31,664,826 B 68,800 rows | 1.6 : 1 | Keeps 53 columns for diagnosis. The ratio is against the decoded color planes, not the original raw file. |
| Walk receipt, M 66 stack | n/a | 271,383 B | n/a | One stage's JSON receipt. JSON is readable but not small. |
| Whole run directory, M 66 stack, early run | 49,772,160 B | 403 MB | 0.12 : 1 | Everything: decoded frame, planes, 105.7 MB of detections, an 18.4 MB preview. Taken 2026-09-14 with 7 of 11 stages run. The working files are 8 times the frame, and they are not the receipt. |
| Arrow buffer compression (LZ4 or Zstandard) | n/a | n/a | n/a | Not used. Arrow IPC supports it, but the engine writes uncompressed files today. Section 14 proposes turning it on. |
The ratio depends on the field. A receipt grows with the number of stars. A frame does not. On a Bayer frame each star costs about 590 bytes in the bank, so the ratio is roughly the frame's size divided by stars × 590. On an S30 Pro sub the two meet near 28,000 stars (16,594,560 ÷ 586.6). A rich Milky Way field could get there. A sparse galaxy field compresses far more than 7.5 to 1.
The slider runs from 200 to 40,000 stars on a log scale; the M 31 sub had 3,760. The bar is the deposit as a share of the sub's bytes, and it fills near 28,000 stars.
Nothing here needs SkyCruncher's code. The tables are standard Arrow IPC files and the receipts are plain JSON. The column names below are real; the paths follow the run directory layout.
# pip install pyarrow import json import pyarrow.ipc as ipc stars = ipc.open_file("run/photometry/photometry.arrow").read_all() print(json.loads(stars.schema.metadata[b"skycruncher"])) # table name, version, word lists measured = stars.filter(stars["status"] == "measured") print(measured.select(["source_id", "band", "psf_flux", "psf_flux_sigma", "snr"]))
# DuckDB queries an in-memory Arrow table by its variable name
import duckdb, pyarrow.ipc as ipc
stars = ipc.open_file("run/photometry/photometry.arrow").read_all()
duckdb.sql("""
SELECT band, count(*) AS n, median(snr) AS median_snr
FROM stars
WHERE status = 'measured'
GROUP BY band ORDER BY band
""").show()
// npm install apache-arrow
import { tableFromIPC } from "apache-arrow";
import { readFileSync } from "node:fs";
const deposit = tableFromIPC(readFileSync("bank/frames/000000.arrow"));
const flux = deposit.getChild("flux_raw");
const sigma = deposit.getChild("flux_sigma");
console.log(deposit.numRows, flux.get(0), sigma.get(0));
// arrow-ipc, the same reader the engine uses
use arrow_ipc::reader::FileReader;
use std::fs::File;
let reader = FileReader::try_new(File::open("run/walk/confirmed.arrow")?, None)?;
for batch in reader {
let batch = batch?;
let ra = batch.column_by_name("ra_deg").unwrap();
println!("{} rows, first ra_deg column has {} values", batch.num_rows(), ra.len());
}
# the run receipt names every product with its hash and size
import json
r = json.load(open("run/run_receipt.json"))
for stage in r["stages"]:
for p in stage["products"]:
print(stage["stage"], p["path"], p["table"], p["schema_version"], p["bytes"], p["sha256"][:12])
Engine tables keep their name, version and settings as one sorted JSON object under the metadata key skycruncher. Bank files carry the word lists that decode their one-byte codes (bands, status words, reasons, builds) in their own metadata, so a reader turns codes back into words without SkyCruncher's code.
A receipt is a measurement, not a copy. A frame is what is on the sky (stars and a galaxy) pushed through the camera's optics and the night's air, plus the glow of the air itself, plus noise. Each part can be held a different way: as rows, as a few numbers per frame, as a model kept once, or not at all. Pick a term to see it alone, what it costs, and which level holds it.
The left picture is a real 10-second sub of M 31. The right picture was drawn without any of its pixels, from the receipt the Layers tab proposes: every Gaia star at its catalog position with the brightness this frame measured, drawn through this frame's optics and air, on its sky glow, with noise at the predicted level. The bottom picture is the real frame minus the drawing (without its noise), so whatever is left there is what the receipt does not hold.
The three choices add a galaxy layer. Galaxy model is M 31, M 32 and M 110 as the night's other 351 subs saw them, with the stars removed: an object model held once, pushed through this frame's optics and air, and only its brightness is new in this frame (16 bytes). Survey projection takes the DSS2 survey plates of the same sky instead, projects them through our camera, and scales them with two numbers per color plane: the version every camera could share today, but the plates are saturated in M 31's bulge. Either way the galaxy is held once, in one shared place, and is not baked into the frame.
Real frame
Synthetic · stars only
RealSynthetic
Real minus synthetic · stars only| Drawn | Per frame | Held once | Leftover structureover the galaxy, against blank sky | Galaxy light |
|---|---|---|---|---|
| Stars onlyGaia stars, this frame's brightnesses, through its optics and air | 52.9 kBbrightnesses, the air, 3 ring numbers per bright star | Gaia DR3+ 677 kB star shape, per camera | 15.9× | 0% |
| + Galaxy modelM 31, M 32, M 110 from the night's other 351 subs | +16 Bits brightness | 481 kBthe object model | 1.12× | 96% |
| a smooth model insteadbulge and disk fit (the first version of this figure) | +160 B | 200 B | 1.6× | 88% |
| + Survey projectionDSS2 red and blue, through our camera | +32 B8 numbers | 47.9 MBtwo survey plates | 7.0× | 104% |
| Galaxy pixels insteadevery galaxy pixel, 4 bands, 16-bit | +2.29 MB | nothing | none | 100% |
The first version of this figure left a dozen bright stars at full strength in the difference. The stars were in the receipt, at the right positions. Their fluxes were not: in some bands missing or negative, in others far too high, and that drawing skipped any star with an impossible flux in one band.
In those rows the joint fit placed the star in one group with an uncataloged detection sitting on top of it: the two profiles overlap by 0.99 at the median, against 0.37 for the bright rows that came out fine, so the fit cannot tell them apart and splits the light arbitrarily. Saturation was not the cause: the frame reaches its 65,535 ceiling on only 37 pixels.
For the rollup this matters most. A design that stores only what the model fails to predict would keep these broken rows as "surprises", while the brightest, most visible stars would have no usable flux. The receipt should carry bright stars on purpose: a flag, the catalog position and magnitude, and a model of the star's core and wings (with saturation where it applies), so a broken fit falls back to the catalog instead of to nothing.
If you keep only receipts, here is what you can and cannot get back.
Raw frames keep everything and cost the most. A receipt keeps the part of a frame that a star survey needs, in a layout built for asking questions across thousands of frames. The trade only works if the measurement is good, which is why every row carries its error bar, its evidence and the build that made it.
One more limit, stated in the code itself: the first community bank's error bars were 20 to 50 times smaller than the scatter the points actually showed, because only one term went into them. The bank now names which terms are in each frame's bars, so a reader can tell a small bar from an incomplete one. Part two depends on getting this right.
A design for storing far less per frame once the models are good enough to predict most of what a frame will show.
Each level stores what the level above it cannot predict. A frame becomes what the session model predicts, plus whatever surprised it. A session becomes what the hardware profile and the star database predict, plus whatever surprised them. Once the models are good, nearly every star on nearly every frame is unsurprising. Then you keep a few running totals per star instead of a row per measurement, and full detail only for the surprises: variables, transients, artifacts. Galaxies and nebulae follow the same rule through object models, which sit beside the star database (section 12).
Today's bank. Every star, every band, every frame gets a row, about 590 bytes per star on a Seestar frame. Everything is kept, including the noise.
One night, one rig, one target: the same ~3,000 stars in ~300 subs. Most frame-to-frame change is shared by every star on a frame: transparency and zero point, sky background, seeing, airmass, tracking drift and field rotation. Each frame becomes a handful of nuisance parameters. Each star becomes one summary: mean brightness, scatter, number of frames, goodness of fit. A constant star's 300 rows collapse to one. A star whose wiggle exceeds what the noise model predicts keeps its per-frame rows, because that is signal.
One camera unit across all its nights. Gain, read noise, dark current against temperature, the flat field and vignetting, color response against a standard system, star shape across the field, nonlinearity and hot pixels do not change between nights. They live in the hardware profile, one per unit: an S30 Pro's telephoto and wide cameras are separate units. Once a sensor's spatial behavior is known, the session model needs fewer free numbers, so session summaries get smaller and more accurate. Contributed darks, bias frames and flats are the fastest way to build one. Section 14 lists these terms in parametric form.
One row per star, in a standard system: position and motion (mostly from Gaia), brightness in standard bands, color, and a description of how it varies (period, amplitude) where it does. The predicted measurement is a function of the star's values, the hardware profile, the session conditions, and where the star landed on the sensor.
One model per galaxy or nebula per band: a map sharper than any Seestar frame, or a fitted profile with a map for what it misses. It is projected into each frame through the same hardware and session profiles as the stars, so a frame's extended light costs a few numbers instead of pixels. Section 12 has the detail and what the M 31 work showed.
For every star on every frame, the levels above make a prediction: a value and the noise it should have. If the observed value matches within that noise, nothing is stored for this frame. The star's running totals are updated instead. If it does not match, the measurement is stored in full as an exception.
Bright stars need their own path. A failed fit is not a surprise about the sky, so a bright star whose flux comes back missing, negative or far from its catalog magnitude should be predicted from the catalog and its saturation model, and the row flagged, rather than stored as an exception (section 06 shows why).
Per star, per band, per hardware unit, the design keeps three sufficient statistics: the count, the weighted sum of residuals, and the weighted sum of squared residuals. They rebuild the mean and the scatter exactly. They also add: session totals add into hardware totals, which add into the star's global totals. The rollup is just addition, and it can be redone in any order.
The same normalization as today, carried one level further. Facts live at the level they belong to, and everything below points up by key.
The noise model is everything. "Within the predicted noise" only means something if the error bars are right. Today's bank shows the danger: the first community bank's error bars came out 20 to 50 times too small. Too small, and everything looks surprising and nothing compresses. Too large, and real variability is averaged away.
A receipt today keeps the stars and drops the light between them. The same predict-and-store rule can bring that light back. Model each galaxy or nebula once, as an object model. Project it into every frame through the same hardware and session profiles that predict the stars. Then store only what the projection could not predict.
Tried on a real frame. The figure in section 06 does this on one M 31 sub. An object model built once from the night's other 351 subs (481 kB held once, 16 bytes per frame) brings back 96% of the galaxy light and leaves structure 1.12 times blank-sky noise. A smooth bulge-and-disk fit in its place leaves 1.6 times (1.48 with GALFIT, 1.33 with M 31's bar). The DSS2 survey plates projected through the frame (47.9 MB held once, 8 numbers per frame) take out the disk but leave the bulge, where the plates are saturated (7.0 times). Storing the same galaxy as pixels would cost 2.29 MB in every frame. One shared reference, held in one central library, plus a few bytes per frame replaces the galaxy's pixels in every frame that sees it: the object model pays for itself in the first frame, and the survey plates after about 21 frames of M 31.
| Where | What | Why |
|---|---|---|
| Once per object and band | The object model and its version, a few MB per map | Every frame of that object reuses it. |
| Measured on M 31 | Object model: 481 kB once, 16 B a frame. Survey plates: 47.9 MB once, 32 B a frame. A smooth fitted model: 200 B once, 160 B a frame. Galaxy pixels: 2.29 MB a frame (one 800 × 600 view, 4 bands, 16-bit) | From the section 06 figure. Pixels grow with every frame; a reference does not. |
| Per frame | Nothing new | PSF, zero point, sky plane and registration offset are already frame parameters. |
| Per session | One stacked residual map over the object's footprint, per band, full or reduced resolution | Model error is the same in every frame, so stacking a night's residuals shows it while the noise averages down. A large residual map is the signal to update the model. |
| Exceptions | Anything that changes during the night: a supernova, a comet, a satellite | Stored like any other surprise. |
| Cost | At the section 13 defaults, about 16 GB a year plus 8 GB once | Under a tenth of either rollup. Apart from raw FITS, it is the only way the archive keeps galaxy and nebula light. |
| Lost | Per-frame pixel noise in the galaxy, and detail finer than the model holds | A new way of measuring the galaxy needs the model and residual maps, or the raw frames. |
From the effort to recover Hubble's Cepheid V1 in Andromeda. Flux is in f19, the light of a G = 19 star.
In the engine, nothing is built yet. No galaxy or nebula model exists in the engine today. The session model's design already describes the prediction this section needs: transparency times the rendered scene, plus sky, with conditions as "a few numbers per moment, not a picture".
Start with what worked: a Pan-STARRS or Hubble template, or a sky-fixed layer fitted from the session's own subs, plus catalog stars, through each frame's PSF, zero point and sky. The research version ran at about 0.2 s per sub on 9.6′ cutouts around V1; full frames will cost more.
Stack each session's residuals over the object's footprint and keep one map per band. It is the extended-light version of an exception, and the evidence for the next model version.
The M 31 experiments showed the Moffat shape misreads S30 frames. The hardware profile holds the shape across the field; the session adds the night's seeing.
Every session that sees M 31 improves one versioned model, and its residual map becomes evidence for the next version. First: model versioning, and a way to merge residual maps from rigs with different PSFs and bands into one correction.
Fit one sharper model to all sessions at once, each through its own PSF, to resolve detail no single Seestar frame shows. First: per-frame PSFs and registration good to a fraction of a pixel, held-out checks that the new detail is real, and the compute to fit it.
A bulge, a disk and a Sérsic profile with a map for what they miss would be smaller and would not depend on outside surveys. First: the experiments so far reach parity at best, cannot separate bulge from disk, and have no dust term, so a model with dust and a single disk component has to beat a survey template on held-out frames. Section 14 lists the same models as a reach.
Emission nebulae as line maps with physical line ratios, dark nebulae as absorption in front of the stars. First: an absolute anchor such as IPHAS H-alpha or our own deep stack, because one frame cannot separate sky from nebula, and a basis that covers scales from about 50 px up to the whole frame.
If exceptions run to a few percent of rows, the cost of a star on a frame drops from about 590 bytes to a few bytes of totals plus the occasional exception row. That is plausibly one to two more orders of magnitude beyond today's 7.5 to 1. The real figure depends on how often the model is surprised. The chart below compares four ways to keep a growing community's data over ten years. Both rollups include the extended-light layer of section 12 unless you switch it off below.
Change any number; the chart, table and arithmetic update.
Log scale: each gridline is ten times the one below. Hover or tap for values.
| Strategy | 1 year | 3 years | 5 years | 10 years | Smaller than raw |
|---|
CanEverything, including measurement methods nobody has written yet.
CannotNothing is lost. The cost is the bytes.
CanAny star-level question: positions, brightness, error bars and evidence on every frame.
CannotAnything that needs pixels: extended light, faint sources nobody looked for, a new way of measuring.
CanEach star's mean, scatter and variability per session, plus every exception in full. With extended light on, galaxy and nebula light as object models plus a residual map per session.
CannotA constant star's value on one particular frame. With no audit sample, there is nothing to re-test a new model on.
CanStandard-band brightness and variability for every star across all rigs, every exception, galaxy and nebula light as object models plus residual maps, and whatever the audit sample allows.
CannotThe exact per-frame noise of each constant star outside the audit sample, detail finer than the object models hold, or pixels.
Every frame SkyCruncher touches gets a fingerprint: a SHA-256 hash, a 64-character code computed from the bytes. Change one bit and the code changes completely, and nobody knows how to make two different inputs share one. The engine, the upload page, the upload server and our research tools all use this same hash. This tab follows one frame's fingerprints from the observer's disk to its receipt and the bank. Then it proposes how the rollup levels could cite the fingerprints of what they summarize, so any number in the star database traces back to the frames it came from.
A FITS file is a chain of HDUs (header and data units). Each is a text header of 80-character cards, then the data, padded with zeros to a whole number of 2,880-byte blocks. Three hashes cover three different slices of the file, and each answers a different question. A fourth, in the engine's stacker, covers the decoded pixels.
| Fingerprint | What it covers | Survives the site fuzz | What it is for |
|---|---|---|---|
| Frame IDthe engine's fingerprint | Every byte of the file: headers, data and padding. | No | The engine's key for a frame: the run receipt, the run ID, the bank's frame table, and duplicates within one batch. |
| data_sha256the upload's data hash | Every HDU's data bytes in file order, each cut to the length its cards declare (BITPIX, NAXISn, PCOUNT, GCOUNT). Padding and headers excluded. | Yes | The upload page's duplicate check, and the upload server's index across batches. Computed on both sides and compared. |
| Quick keyidentity without the pixels | Nine identity cards exactly as written (INSTRUME, TELESCOP, DATE-OBS, EXPTIME, EXPOSURE, BITPIX, NAXIS1 to 3), a newline, and the first 2,880 data bytes. | Yes | Asking the server "do you already have this frame?" before reading it off a slow drive. The server answers yes or no, never where or whose. |
| Plane hashin the engine's stacker | The decoded color planes: their names, geometry and samples, after a version tag. | Yes | Catching a header-edited copy of a frame the stack already kept, which the frame ID cannot see. |
Why more than one? Reading every byte off a Seestar over USB runs at about 16 MB/s. The quick key lets the upload page skip frames the server already holds after reading one header and one block. It is an identity check, not an integrity check: a change deep in the pixels leaves it alone. data_sha256 is the integrity check. The page computes it while the frame uploads, and the server computes it again from the stored copy, so the page's claim is never the only record.
Pick a FITS file, or let the page draw a small synthetic one. The page computes all three fingerprints with exactly the upload page's definitions. Then change the frame and watch which fingerprints notice.
Your file never leaves this browser. The page reads it (in 4 MiB pieces when it is over 64 MiB) and hashes it on your device with the browser's own SHA-256. The demo makes no network requests and saves nothing; the page's only outside request is the site's analytics beacon, which never sees your file.
…What the engine would key this frame by.……Each change is built the way the real fuzz builds an upload: a new header, then slices of the original file. Nothing is copied to disk.
| Version | Frame ID | data_sha256 | Quick key |
|---|
The gap. The engine's frame ID covers the whole file, so the privacy fuzz gives an uploaded frame a new ID. To the engine, the observer's original and the uploaded copy are two frames, and the same frame run from both gets two run IDs. The decision that a frame is keyed by its bytes (2026-09-13) names this exact case as the time to revisit it: "the same observation carries two hashes." data_sha256 is the bridge, and the receipt does not carry it yet. That is the first recommendation below.
A chain of custody means each step records what it received, by fingerprint, so anyone can check that nothing was swapped or altered on the way. The top row is the path of one uploaded frame today. The bottom row is how the chain could continue up the rollup levels from Part two.
On the observer's disk. The upload page reads the header and the first data block to make the quick key, then asks the server which frames it already holds and skips those. It computes data_sha256 when it reads a frame to send it. Frames with more than 512 MiB of data are not hashed.
In the browser. The site cards move by a stored random offset of up to 5 km. Cards computed from the site (altitude, azimuth, airmass, hour angle and the like) are recomputed for the moved site or removed, and two HISTORY cards mark the file. The upload is built as a new header plus slices of the original file, so the data bytes are never read, copied or changed. None of the quick key's nine cards is a site card. So the quick key and data_sha256 come through unchanged, and the whole-file frame ID does not.
The upload server (a Cloudflare Worker, a small program that runs at Cloudflare's edge) stores the file, then hashes the stored copy's data units itself, streaming them through a digest. It writes a small hash record beside the batch and keeps two indexes in its key-value store: data_sha256 to the first stored copy, and quick key to data_sha256. A frame that is already stored is flagged, never refused or deleted. The batch manifest lists, file by file, whether the page's data_sha256 matched the server's.
The engine fingerprints the file it was handed: for an upload, the fuzzed copy. The run ID is a SHA-256 over the build, the runtime switches, each input as its role and frame ID, and the stage list, so the same request twice is one run. Every product is listed with its SHA-256 and byte count. The solve, walk, photometry and stack stages add a decision digest: a SHA-256 over what the stage decided, with every number rounded to the grain the measurement can resolve, so two machines that agree within that grain write the same digest.
Designed but not yet built: digests for preprocess, bands, detect, noise model and grade. Until they exist, a full run's own rolled digest is empty. The receipt does not carry data_sha256 or the quick key.
One file per frame that is never edited, a frame row keyed by the frame ID, and every row stamped with the build that measured it. A bank's identity is the rows it holds, never the file's bytes or their hash (decided 2026-09-15), because the bytes also depend on the Arrow library's version and settings. So a digest that cites a deposit should cover its rows, or reuse the photometry decision digest, which already identifies the measurement.
The session profile (designed in Part two) would cite a Merkle root over its frames, one leaf per frame: its frame ID, data_sha256, run ID and photometry decision digest. The stacker already keeps exactly this kind of leaf list: its table of inputs names every frame offered, kept or excluded, with its frame ID and plane hash.
Would cite the roots of the sessions it was fitted on. Today a run names its hardware profile by the SHA-256 of the profile table it wrote, which moves whenever the measured profile does.
Each release would carry one root over the totals and sessions it was built from, and each star's row the roots of the sessions that fed its running totals. Because totals add up, anyone holding those session summaries can add them again and check the star's row.
Each model would cite the SHA-256 of the survey images it was drawn from. Our research tools already record the address, size and SHA-256 of every survey image they download, including the DSS2 and Pan-STARRS cutouts behind the real-against-synthetic figure.
A Merkle root is one hash that stands for a whole list. Hash each item into a leaf, hash the leaves in pairs, then the pairs in pairs, until one hash is left. Change any item and the root changes. To prove one frame is in a session, show the frame's leaf and the sibling hash at each level on the way up: about 20 hashes for a million frames, 640 bytes. So a rollup does not have to keep its frames' rows to stay checkable. It keeps the root, and the list of leaves sits beside it.
Two habits keep it honest, and both have precedent. Tag what is hashed: the engine's plane hash starts with its own name and version. And mark leaves and inner nodes differently, as the public Certificate Transparency logs do (RFC 6962), so a leaf can never pass for a node.
// one leaf per frame in a session; the same shape one level up for sessions, profiles and releases leaf(frame) = sha256( 0x00 || "skycruncher.session.leaf.1\n" || frame_id || data_sha256 || run_id || photometry_decision_digest ) node(l, r) = sha256( 0x01 || l || r ) root : sort the leaves by frame_id, hash them in pairs left to right, let an odd one out move up unchanged, repeat until one is left proof : the sibling hash at each level, about log2(n) of them // the row stores inputs_root, n_inputs and root_rule (this recipe in words, so the string is never an oracle)
Beyond the frame's own three, these are the fingerprints already written today, and where.
| Hash | What it covers, and why | Status |
|---|---|---|
| In the engine | ||
| Product hash | SHA-256 and byte count of every table and JSON file a stage writes, listed on the run receipt. It is what makes a row citable: the row names the bytes it was read from. | Built |
| Decision digest | What a stage decided, every number rounded to the grain the measurement can resolve, so two machines agree. The run rolls its stages' digests, in order, into one. | Builtsolve, walk, photometry, stackDesignedthe other five stages |
| Host digest | Which build decided it: version strings and other facts about the machine, kept out of the decision digest on purpose. | Built |
| Run ID | SHA-256 over the build ID, the runtime switches, each input as role and frame ID, and the stage list. It doubles as the job ticket, and every ingredient is written beside it so anyone can recompute it. | Built |
| Build ID | SHA-256 over the commit, build switches, toolchain and target, plus the SHA-256 of the binary that actually ran. | Built |
| Hardware profile ID | For now, the SHA-256 of the hardware profile table a run wrote, until the profiler mints a key of its own. | Built |
| Rig hash | A salted hash of the telescope, body and lens serials, so records shared beyond the observer's machine can name the rig without carrying its serials. The code that mints it exists, and every receipt still says "not carried". | Designed |
| Reference releases | Catalog and index releases are published with a manifest of byte sizes and SHA-256, and verified as they land. | Built |
| In the research tools | ||
| Dataset manifest | SHA-256 of every file in a dataset, and one digest over the sorted lines "path, bytes, sha256": a one-level Merkle list. Data enters only through a hash-verified landing, and copies are verified against it again later. | Built |
| FREEZE manifest | SHA-256 of the engine binary, the index manifest and an experiment's scripts, written before the experiment looks at results, so a later reader knows exactly what ran. | Built |
| Survey image manifest | Address, byte count and SHA-256 of every downloaded survey image. | Built |
Proposals, not decisions. Each one is small on its own.
data_sha256, quick_key and a one-line data_sha256_rule to the run receipt's frame block, as a minor version bump. The engine can hash the data units in the same streaming pass that already makes the frame ID. This joins a receipt to its upload records, and to the observer's original file across the fuzz.header_edits field naming the fuzz marker cards the frame carries, such as the privacy rule's version, so a reader knows why the frame ID will not match the observer's own file.data_sha256, run_id and photometry_decision_digest columns to the bank's frame table, as a minor bump. Then a frame deposited twice, once from the original and once from the fuzzed copy, is caught, and every bank row leads back to its receipt.inputs_root, n_inputs and root_rule, with the list of leaves stored beside them, using the recipe above. The stack is the cheapest place to start, because its inputs table already holds the leaves.The demo's hashing mirrors the upload page's own code line for line, and a test checks it against that code: on synthetic frames, a real Seestar sub, a multi-table index file, and frames passed through the real site fuzz.
One direction goes further than receipts: archive frames as forward-model parameters plus a residual quantized at the noise floor. Named model components compete to explain each pixel, and each one is admitted only if it earns its place. Storage becomes minimum description length: the cost of the model's parameters plus the cost of whatever the model leaves unexplained. That is the same rule as Part two, applied to pixels instead of star measurements. These are the ideas that bear on this page.
Stabilize the noise (an Anscombe transform), round each pixel to about half its noise, then entropy-code the integers (Rice or Zstandard). Rounding to half a sigma adds about 1% to the noise (√(1 + 1/48) ≈ 1.01). Most of the gain needs only a noise model, not a model of the scene. One estimate is 4 to 8 MB for an 18.8-megapixel frame. At that bit rate an S30 Pro sub would be about 1.8 to 3.5 MB instead of 16.6 MB, so a 30-day raw window at the section 13 defaults would hold about 350 to 700 GB instead of 3.3 TB. This is untested on Seestar subs; benchmark against FITS fpack first.
With the codec, pixels for 1 frame in 50 would cost about 84 to 168 GB a year at the defaults, close to the 106 GB the audit rows cost. Pixels let a new measurement method be re-run on the audit sample, which rows cannot.
Zstandard suits parameter tables, and LZ4 fast staging. The engine writes uncompressed Arrow today. One-byte codes and sorted keys should shrink well; 64-bit floats less so. Measure on the M 31 deposit before deciding.
A radial vignetting polynomial, dark current that scales with exposure and temperature, a bias pedestal, hot pixels found by staying put across dithered subs, column offsets, nonlinearity and a saturation level. Each is a few numbers or a sparse list instead of a full-frame map, built from contributed darks, bias frames and flats.
A low-order sky polynomial, transparency (zero point, optionally a coarse cloud grid), airmass extinction, the PSF across the field, tracking drift, and a per-band registration offset. This is what fills the 200 bytes a frame assumed in section 13.
Tag each exception with what caused it. A satellite or aircraft trail becomes one row of endpoints and width instead of a row for every star it crosses. A cosmic ray is a single-sub hit. Known asteroids, checked against the Minor Planet Center's orbits, stop false variable-star claims.
Cross-match the star database with AAVSO's VSX catalog for known variables. A star counts as variable only when its scatter beats matched constant stars in the same session, which is the totals check of section 09.
Outside catalogs may pull a value toward zero, never put light where none is measured. A new model term must improve predictions on held-out data. And all modeling stays on the raw Bayer planes, because demosaicing correlates neighboring pixels. Section 11 adopts all three.
Keep only model parameters, about 50 to 350 KB a frame by one estimate, and drop the residual. This is high risk: it erases uncataloged transients and asymmetric structure. First: a model that explains every component of a frame, proven on held-out frames, with the residual kept somewhere for a while.
Split each frame into tiles: empty sky as formulas only, star fields as model plus quantized residual, nebula cores and anomalies at full resolution. First: the codec above, and a scene model good enough that empty-sky tiles rebuild from formulas within the noise.
Sérsic galaxies, regularized grids, galactic cirrus scaled to Planck dust maps, emission and reflection nebulae. Section 12 covers this as object models. First: each model must beat a survey template on held-out frames.
Moonlight, twilight, city light domes, airglow and zodiacal light, each pinned by an ephemeris or a light-pollution map, instead of an anonymous sky polynomial. First: evidence that the polynomial is not enough for star photometry, and local copies of the outside datasets.
An observation plan could tell each zone of a frame how to be processed: keep raw pixels in a window around a transient candidate and summarize everywhere else. First: the codec and tile split above, and a way for the upload page to carry the rule.
Transit and eclipsing-binary templates, or a free amplitude per exposure, in the star database's variability description. First: enough nights per star, and template fits checked on known systems.
Export each star's parameter uncertainties and their correlations from the solver. First: the joint fit of hardware and stars from section 11 has to exist.
Formats, schemas and byte counts are read from SkyCruncher's engine and its measurement reports as of October 2026; the code is not public yet. The M 31 frame in Part one is the author's own: one 10 s sub from a ZWO Seestar S30 Pro on the night of 2026-09-18, with the night's other 351 subs for the object model. The M 31 findings in section 12 also use data from the Seestar Collaboration's M 31 project, with permission.