The descent
You are one of six million objects in the vault. No teacher has ever found you. Scrolling this page is the journey that finally puts you in front of a class.
you are here: 1 of 6,000,000 · unfound
the problem
A K-12 teacher does not have a content shortage, they have a sifting problem. The current Learning Lab exposes over 6 million objects and 50,000+ collections, yet a typical educator actively keeps only about 5 supplemental plus 2 core materials in rotation (requirements doc). Districts touch an average of 2,982 distinct edtech tools a year, the scan found 64% of user-generated resources are "not worth using," and the incumbent was the slowest platform reviewed, where the word "collection" carries no schema a teacher recognizes. Around 30% of ELA and elementary social-studies teachers start with a Google search at least weekly, and out-of-field teachers (about 25% churn into a new grade or subject each year) get no point-of-use background to teach safely. The job is not "more," it is "just right for me."
the vault
Every catalog lists what the current Learning Lab exposes: 6M+ objects across 50,000+ collections. But a time-strapped teacher keeps only about five supplemental and two core materials at a time. The volume that looks like richness is, to them, the reason you stay buried, and 64% of what they find elsewhere is judged not worth using.
the readiness gate
So discovery never searches the whole vault. A readiness gate sorts the millions of EDAN / Open Access records into transform-now, queue-for-enrichment and defer, and only objects that can become classroom-ready cross into a small curated set. The pile never reaches the teacher by construction. This is the bet: just right, not everything.
the trust ladder
Now you are made trustworthy, and every claim about you is earned from evidence, never asserted. Standards alignment is confirmed, not sprayed across fourteen tags. Every generated line is cited back to your own record, or it is parked. Accessibility is checked. Then a human curator says yes. Only then do you wear the green badge, the one signal reserved for trust.
one source
You are described once, in a frozen contract, and from that single source we derive the teacher screen, the open API, and the assistant tools. There is no second story to drift out of sync. The AI assistant is a customer of the same truth the teacher sees: it accelerates the work, it is never the only way to do it.
the action layer
A collection has no shape a teacher recognizes. So you are shaped into the formats they actually teach from: a bellringer, a hook, an inquiry routine of observe, reflect and question, an activity, an assessment, an exit ticket, plus leveled text and slides. Each one is generated from your record and cited, so an out-of-field teacher can lead it without fear of getting you wrong.
for every learner
The students who most need a supplement, multilingual learners and striving readers, are the ones the market serves worst. So you carry leveled tracks and full accessibility: the same object, met at each reading level, with alt text, plain language and a 508 transcript. 86% of teachers adapt what they use, so here adaptation is built in, not bolted on.
the last mile
Found, trusted, shaped, leveled. The last mile is the room itself: a teacher view and a separate student view, a whole-class present mode, and one move into the tools teachers already use. The teacher view runs today; the student view, present mode and deep LMS integration are the mile we are building now. We mark what is built and what is next, honestly.
send anywhere
A lesson is only as useful as the classrooms it reaches, so the strategy is connectivity by open standards, not lock-in: a teacher builds a trusted object-lesson here, then takes it into the LMS they already use - no migration, no walled garden (the incumbent's actual weakness). Three connectors ship today: an embeddable widget for any page, a Google Classroom share, and standards export to CSV and Markdown. The roadmap upgrades those to standards-grade reach - one certified LTI 1.3 + Deep Linking tool that places a lesson inside Canvas, Schoology, Brightspace, Moodle and Blackboard at once (one build, every LMS, which is the whole point), a deeper Google Classroom path since it is the one major platform that speaks no LTI, an oEmbed endpoint so a link auto-embeds across the open web, and CASE for machine-readable standards in and out. Interoperability is not the headline - the trusted Smithsonian content is - but it is the moat that lets that content win the room. Authentication and rostering are out of scope this version, so identity-dependent pieces (roster pull and grade pass-back via OneRoster and LTI AGS) are a stated open-standards roadmap, not a claim.
pain to solution
Learning Lab Studio is aspirin, not a vitamin: each named teacher pain points at a capability that cures it. The current Learning Lab exposes over 6 million objects and 50,000 collections (00_overview.md), yet a typical educator actively keeps only about 5 supplemental plus 2 core materials, so the answer is a curated, gated set, not everything. Five of these pairs are shipped and exercised end to end (14_demonstrated_capability.md): the green badge publishes an alignment only when its code resolves to a real 1EdTech CASE indicator, and the faithfulness gate routes any draft below 0.90 to human review (05_ai_assisted_generation_stack.md). Three pairs, marked with dashed connectors, are the honest next mile: whole-class delivery, crawlable LRMI discovery pages, and the student view. The figure reads left to right: the pain, then the move that answers it.
ai readiness
We landed and scored the full EDAN / Smithsonian Open Access corpus: 14,242,091 records across 29 units, every number read live from the same rubric the factory runs (ai_source_readiness doc). Rights are the one solved problem - 100% CC0, the licensing risk that sinks most supplemental content is simply gone. But media is the binding constraint: only 36.8% of records carry a usable image, which is exactly why 63.2% fall into queue-for-enrichment and just 34.5% can transform now. The evidence that curation beats indexing is the education signal at 0.7% - indexing the raw pile would bury a teacher under 99.3% noise. This is MEASURED for EDAN / Open Access (14.2M) only; the incumbent Learning Lab surfaces about 6M, but we have not ingested it, so the overlap between the two is unknown until the content audit (blocked, post-award), and standards coverage is partial.
the data quality check
Before claiming anything is AI ready, we ran the data-quality checks live over all 14,242,091 records. Integrity is pristine: zero duplicate IDs, 100% CC0 rights, 100% public, and zero contradictions between the media flag and the media count (every record that has media has CC0 media). So the work is not cleaning dirty data, it is closing a coverage gap. Field completeness is uneven: title, link and rights are 100%, topic 89%, place 87%, names 85%, date 78%, but usable media is only 37% and object type only 26% - the two shortest bars, and the real lever. And the gap is not random: it clusters by unit. Imaging runs 95% at the National Portrait Gallery and 88% in Birds but 4% in invertebrates and near zero in anthropology; the education signal lives in NMAAHC (19%) and the Portrait Gallery (15%) and is near zero in the giant natural-history units. The readiness score is bimodal, peaking at 7 and 12, and the gap between the two peaks is exactly the five media points - so media alone is the swing factor. That is why curation targets the units where readiness already lives rather than boiling the ocean.
the data we stand on
Here is the whole data picture, honestly. One external SOURCE flows live today: EDAN, the Smithsonian Open Access bulk - 14,242,091 CC0 records across 29 units - the only source we ingest, and one we never own (the system of record stays at the Smithsonian). From it we DERIVE three stores we own: a faithful bronze landing plus a 9-factor readiness score in BigQuery, an append-only transformation registry that carries a W3C PROV chain of custody on every step, and a consent-gated usage-signals store that stays empty until a teacher opts in. The set a teacher actually SEES is a human-gated curated slice of finished lessons, today a working set of 232 records rather than the raw pile. Standards are a real 17-indicator seed across three frameworks (CCSS ELA, NGSS, C3) with a 1EdTech CASE parser ready to load full frameworks. The legacy Learning Lab corpus - 6M+ objects, 50,000+ collections, 600,000+ users - is the content-audit target the RFP names, blocked until a post-award data agreement, and it rides the identical pipeline the day a feed exists.
the life of a record
Every record travels the same path, and four of the eight stops are gates that can stop it. A raw EDAN line is ingested losslessly into one typed SourceAsset, landed to a resumable bronze layer, then scored: a readiness gate sends only classroom-worthy records into the factory. Inside the factory, generation is grounded only in the catalog record with citations attached from that record, and an automated faithfulness check parks likely hallucinations; standards are model-proposed then format-and-existence validated, so the green badge is earned, not asserted. Then the hard gate: a human curator records a verdict, and publish is a no-op without it. Only published resources are served, and discovery is hard-scoped to that curated set, never the 14.2M. Finally, opt-in usage signals mine search misses into a what-to-build-next backlog that closes the loop. Every stage appends a provenance record, so the lineage from catalog record to lesson is unbroken.
the AI, governed
Ask "where exactly is the AI?" and the platform itself answers. Every AI activity - drafting lesson text, writing alt text, proposing standards, ranking search by meaning, the assistant drafting a quiz a teacher must commit - is enumerated in a machine-readable register the product serves, each entry stating what data goes in, who sees the output, and where the human gate sits. A build gate keeps that register complete: a model call that is not registered fails the build. And the register is governed, not just read: in AI Governance (Policy Studio's sibling) the Smithsonian's own steward can block any activity, and the switch enforces at the code path within seconds - the factory stage skips, search degrades to lexical ranking, the assistant drops to its no-AI path - with every decision attributed in a durable trail. Switch everything off and the product still works; only new AI drafting stops. It is the pattern Smithsonian already trusts internally (SIDECAR: AI drafts metadata, a person reviews), made inspectable and made governable.
where it sits
Pull back. We do not recreate the Smithsonian content or strand the equity it has built. We ingest the objects, the open-access media and the standards, we add the legacy Learning Lab collections once a post-award data path opens, we add this creation-and-trust layer, and feed a modernized delivery surface that keeps the existing users, collections and URLs. Build on the content, replace the platform, migrate the equity, prove each teacher pain relieved.
strategy / our intention
Smithsonian's value is its content, not the incumbent's code. We build on the corpus that already exists - EDAN and Open Access records, the Learning Lab legacy of over 6 million objects and 50,000+ collections (about 2,900 of them authored by Smithsonian educators), the unit education-URL repository, and CASE Network 2 standards - and we leave Smithsonian's systems of record untouched (02_architecture). What we replace is the platform layer that is slow, serves image-of-text, and has no teachable roles. Learning Lab Studio is the creation and trust engine in the middle: it profiles assets, scores readiness (0-14), runs the human-in-the-loop transformation workflow, and feeds a modernized Learning Lab 2.0 delivery surface for teachers and students. We migrate the equity rather than strand it - 600,000+ user accounts, 50,000+ collections triaged keep / transform / archive, and permalinks preserved with redirects and canonical tags (11_proposal_and_open_decisions) - and post-award we assess, then modernize: optimize what is salvageable and replace only what must change, with AI readiness as the gate.
the honest state
This is the honest tally. All 179 requirements distilled from the solicitation and the environmental scan, each mapped to our solution and adversarially verified against the code: an auditor opened the cited file for every claim and downgraded anything a document merely described. 16 are shipped and exercised end to end, 90 are partial (the load-bearing mechanism is built, some acceptance sub-criteria remain), 17 are designed, 30 are honest gaps, and 26 are process commitments that belong in the schedule. The mass is deliberate: the trust-and-content engine that decides what reaches a teacher - readiness gating, the human approval gate, standards verification, provenance - is the most mature, while the surrounding surfaces (a full educator authoring UI, the student view, live LMS integrations, legacy migration) are honestly earlier. We mark what is built, what is designed, and what the engagement funds, and we do not inflate. The riskiest thesis, that a contracts-driven human-gated factory turns a 14.2M-record raw corpus into classroom-worthy, standards-verified, provenance-bearing resources without AI output ever reaching a teacher unreviewed, is the part we have proven end to end on a working slice.
today
This is not a slideshow. 232 Smithsonian objects across ten units are published through this exact path: gated, generated, faithfulness-checked, and human-reviewed. Real records, honest scale, with room to grow.
what holds it together
Trust is built in
The green badge is earned through alignment, faithfulness, accessibility and human review, never declared. Trust is a property of the object, not a marketing line.
AI assists, never replaces
The assistant co-pilots discovery, building and adaptation, but it is never the only path and it never publishes on its own. A human stays in the loop.
Honest by construction
Every generated claim is cited to the source or parked. The scale you see is the real scale: 232 objects today, not a number we wish were true.
Built on what exists
We sit on top of the Smithsonian content and the existing Learning Lab, ingesting and migrating rather than recreating. We replace only the layer that has to change.
One source of truth
The data contract is defined once and derived into every surface, so the teacher screen, the API and the assistant can never drift apart.
Accessible to every learner
Leveled text and accessibility are gates an object must pass, so the students who most need a supplement are served first, not last.
the system, in layers
The whole system in five layers, from the classroom edge down to the contract-bound platform, each tagged built-today versus new. The trust-and-content engine (layers 3-4) is the most mature; the surfaces around it are honestly earlier.
Distribution & access edgenew
Syndicate anywhere by open standards - embeddable widget, Google Classroom share, standards export, answer/search engines. CC0 is what lets us syndicate where no other museum can.
Serve front - the teacher appnew
Understand · Discover · Render · Author. The teacher + student surfaces, over one contract.
The factory - governed productioncuration built · rest new
Readiness gate · Object Lens · grounded generation · standards validation · accessibility · the human curator gate. The governed counter to open AI generation.
Content intelligence & corpusmostly built
EDAN ingest · 9-factor readiness score (BigQuery) · append-only transformation registry (W3C PROV) · consent-gated usage signals.
Platform - contract-bound & portablebuilt
Pydantic contracts as the single source of truth · ports/adapters (hexagonal) · Cloud Run · Batch · Vertex · GCS · BigQuery.
Signals loop: the serve front (2) feeds analytics (4), which feeds what-to-make-next back into the factory (3).
every requirement, mapped
All 179 obligations distilled from the solicitation and the environmental scan, each mapped to a real code path or design doc and adversarially verified - an auditor opened the cited file for every claim and downgraded anything a document merely described. Summarized by domain here; the full per-requirement matrix ships in the proposal. These are capability-level obligations, one per promise the solicitation extracts; at build time each decomposes into roughly five to ten technical features, so the engineering backlog behind them runs past a thousand stories. The product plan organizes them into 22 application modules and 495 named features, every one tracing back here.
| Domain | Total | ● | ◐ | ◇ | ○ | ⚙ |
|---|---|---|---|---|---|---|
| Content strategy & curation | 14 | 2 | 10 | - | 1 | 1 |
| Discovery & search | 12 | 1 | 11 | - | - | - |
| Asset model, authoring & templates | 14 | 3 | 8 | 2 | 1 | - |
| Object-based learning | 12 | 3 | 6 | - | 1 | 2 |
| Distribution & integration | 11 | 1 | 7 | 2 | 1 | - |
| Standards alignment | 9 | 2 | 4 | - | 2 | 1 |
| Teacher & student experience | 15 | 1 | 8 | 4 | 2 | - |
| Accessibility, inclusion & equity | 11 | - | 7 | - | 2 | 2 |
| Governance, review & trust | 12 | 2 | 3 | 1 | 3 | 3 |
| Architecture, data & portability | 12 | - | 6 | 2 | 3 | 1 |
| Security, privacy & compliance | 14 | 1 | 3 | 5 | - | 5 |
| Analytics & measurement | 12 | - | 7 | 1 | 4 | - |
| Migration & continuity | 12 | - | 5 | - | 6 | 1 |
| Process & delivery method | 16 | - | 4 | - | 3 | 9 |
| Promotion & marketing | 3 | - | 1 | - | 1 | 1 |
| Total | 179 | 16 | 90 | 17 | 30 | 26 |
reach
A lesson is only as useful as the rooms it reaches. Three connectors ship today; the roadmap upgrades them to standards-grade reach - one build, every LMS. Identity-dependent pieces are a stated roadmap, since authentication is out of scope this version.
| Channel | Open standard | Status |
|---|---|---|
| Embeddable widget (any page) | iframe | ships today |
| Assign to Google Classroom | CourseWork | ships today |
| Standards export | CSV / Markdown | ships today |
| Place inside any LMS at once | LTI 1.3 + Deep Linking | roadmap |
| Paste-to-embed on the open web | oEmbed | roadmap |
| Standards in + out | CASE 1.1 | roadmap |
| Roster + grade pass-back | OneRoster 1.2 · LTI AGS | roadmap |
the record's path
Every record travels the same path, and four of the eight stops are gates that can stop it - which is how AI output never reaches a teacher unreviewed.
| # | Stop | What happens | Gate |
|---|---|---|---|
| 1 | Ingest | A raw EDAN line becomes one typed SourceAsset, losslessly. | - |
| 2 | Land | Written to a resumable bronze layer (GCS / Parquet). | - |
| 3 | Score | 9-factor readiness routes it: transform-now / queue-for-enrichment / defer. | ● gate |
| 4 | Generate | Grounded only in the record, with citations; a faithfulness check parks likely hallucinations. | ● gate |
| 5 | Validate standards | Model-proposed, then format- and existence-checked - the green badge is earned. | ● gate |
| 6 | Curator review | A human records a verdict; publish is a no-op without it. | ● gate |
| 7 | Publish & serve | Only published resources are served; discovery is hard-scoped to the curated set, never the 14.2M. | - |
| 8 | Signals | Opt-in usage mines search misses into a what-to-build-next backlog - the loop closes. | - |
Learning Lab Studio - the classroom-readiness layer for Smithsonian objects. Companions: the story at /proposal · the two-year plan at /plan.