How Lantern decides
Every video gets read before your child ever sees it. The first reader looks at the title, the channel, and the words the video uses about itself. That reading is fast. It covers everything. A second reader then watches the video itself. It sees the pictures. It hears the words. It checks whether the video is really what its title claimed. Watching costs money. So it runs in the order your child is most likely to see next. Videos already on your child's page get watched first. If the second reader finds something the first one missed, the video comes off every page it is on. You are told why. You can open the wall yourself, any time, and see what made it through. Your report card shows how many of your child's videos have been watched. Want that number to climb faster? Add your own Gemini key in settings. Your child's page then moves to the front of the line.
What the reader finds becomes fourteen small scores. Each one covers a different thing. It might be how scary a video feels, or how loud and busy it is. You set the highest number allowed on each dial. You do this when you set up your child's page. Evening hours can carry a stricter number than daytime. A video above your number on even one dial gets dropped. It does not matter how good the rest of it looks. When a reader cannot tell what a video carries, that dial counts as unsafe. It never counts as safe by default.
You do not need to learn all fourteen names. You can write, in your own words, what worries you about what your child watches. Lantern turns that into the right numbers on the right dials. More than one reader looks at anything that matters. Those readers come from different companies. So one reader's blind spot is never the only opinion your child's page gets. When readers disagree, the video does not quietly slip through. It waits for a person to look at it instead.
Right now, the readers are general-purpose thinkers. We ask them to act like a careful babysitter. The next version is a reader trained just for this job. It will be small enough to run on your own computer. It will be able to grow the way your child grows. It can tighten or loosen as they get older. It will not stay fixed at the day you set it up. Nothing about that reader is finished yet. Nothing here is a promise of when it arrives. We are building it this way for a reason. You should be able to see every dial. You should be able to move it, too, instead of trusting a machine to already know best. Read why in the Primer problem. Or keep reading below for the long version, with its sources.
The rest of this page is the version we wrote for someone who wants to check our sources.
For the curious: the long version
This is the referenced version of Lantern's method, written for a reader who wants to check the sources. Every claim here links to something you can open. Receipts for each link, including the date it was fetched, are in _SOURCES.md.
1. What a dial is
A dial is a rating from 0 to 10 that describes one property of one video. Zero means the property is absent from the video. Ten means the property is extreme for a children's media context, not extreme for media in general. A horror film and the most menacing thing that plausibly reaches a child's feed both sit at the top of it.
The ratings are produced by language models reading what the platform publishes about a video: title, channel, description, tags, duration, and a transcript where the platform provides one. Lantern does not download video files.
A second tier watches. Google's Gemini accepts a public YouTube URL as an input and ingests the video on its own servers, so the model observes frames and audio directly: what is depicted, what is said, the cut rate, the loudness range, and whether the video's own metadata described it honestly. Only the public URL leaves Lantern. No frame, no audio track and no downloaded file ever reaches a Lantern machine, which is what keeps this tier inside the metadata-only ceiling the platform's terms impose. A watch costs roughly seventeen times a metadata read, so the tier runs as a priority queue rather than a pass over the whole library: a video already on a child's shelf is watched before a video that is merely admitted. Where a watched verdict and a metadata verdict disagree, the watched one governs, and a watched refusal removes the video from every shelf carrying it.
Each child has ceilings, one per dial, and evening ceilings may be stricter than daytime ceilings. A shelf is built by taking candidate videos and removing every video whose rating on any dial sits above that child's ceiling for that dial.
The failure direction is deliberate and it is the part worth checking. Where several readers disagree, the strictest reading is the one that gates. Where a rating is missing or cannot be read, the video is not treated as a zero on that dial; unknown is not safe. Readers are instructed to report low confidence rather than score an unobserved property as absent. This matters because the metadata is adversarial in the specific sense Papadamou and colleagues measured: inappropriate content aimed at young children mimics or is derived from content that is appropriate for them, and a child browsing from a benign starting point is likely to reach it.
2. Where the fourteen axes came from
The seed was Common Sense Media's public rating rubric, which is the closest production system to what Lantern needed: nine content categories, each rated 0 to 5 dots, plus an ease of play assessment for games. Lantern departs from it in three ways. The scale is 0 to 10 rather than 0 to 5, because ceilings need finer resolution than review dots. Violence and scariness are separate axes rather than one combined category. And the set is sized to what a reader can actually judge from published metadata.
Each axis then had to earn its place against primary literature. The mapping:
| Axis | What it rates | Primary literature |
|---|---|---|
scariness | threat, menace, dread, horror imagery | Harrison and Cantor 1999; Garrison et al. 2011 |
sensory_load | pacing, loudness, busyness, and density of impossible events | Lillard and Peterson 2011; Lillard et al. 2015; Christakis et al. 2004 |
consumerism | product, brand and purchase pushing | Kunkel 1988; Radesky et al. 2020 |
violence_intensity | amount, graphicness and realism of violence, and whether it is rewarded | AAP 2016, Virtual Violence; Bandura et al. 1963 |
language_profanity | profanity, slurs, insults, trash talk | Coyne et al. 2011 |
sexual_romantic_content | innuendo, sexualised framing, explicit reference | AAP 2010 |
substance_use | alcohol, vaping and drugs, weighted by glamorisation rather than presence | Dal Cin et al. 2013 |
risk_imitation | copyable dangerous stunts and challenges | Bandura et al. 1963; Laghmiche et al. 2024 |
parasocial_influencer_pressure | host bonding used to persuade, including sponsored selling | Bond and Calvert 2014; Coates et al. 2019; De Veirman et al. 2019; Alruwaily et al. 2020 |
sleep_disruption | arousal against wind-down, the evening question | Garrison et al. 2011; Garrison and Christakis 2012 |
misinformation_pseudoscience | false or pseudo-authoritative claims presented as fact | AAP 2016, school-aged children |
body_image_pressure | appearance ideals, diet talk, ranking by looks | Grabe et al. 2008 |
prosocial_modeling | kindness, cooperation and empathy modelled | Coyne et al. 2018; Mares and Woodard 2005 |
educational_value | substantive intentional teaching | Mares and Pan 2013; AAP 2016, young minds |
Two of these carry an explicit evidence caveat. The body image literature was collected mostly in adolescent and adult female samples, so it is extrapolated downward with caution. The misinformation axis rests on policy language about inaccurate and unsafe content rather than on a per-video effect study.
What was deliberately left out is as much of the design as what was kept.
- Desensitisation. Funk et al. 2004 documents it, but it is a cumulative outcome of repeated exposure, not a property of one video. A reader cannot rate it separately from the violence and scariness already rated.
- Representation and diversity. Common Sense Media rates it. Lantern does not, because a 0 to 10 quantity of representation is a values dimension rather than a risk dimension, and it is largely invisible in text metadata. Degrading content is caught by the language and prosocial axes instead.
- Screen time, background television and co-viewing quality. These matter, and Kirkorian et al. 2009 and the AAP's school-age statement show how much, but they are properties of how a family uses media, not of a video.
- Cyberbullying and contact risk. Real platform risks that a parent-curated player without comments, search or messaging does not expose.
- Platform advertising load. Injected at playback and not observable from the video's published metadata. The advertising embedded in the content itself is covered by the consumerism and influencer axes.
3. How the readers are kept honest
Three rules, stated as what they achieve rather than how they are built.
More than one reader looks at anything that matters, and the readers come from different model families. That follows the panel result reported by Verga et al., where a panel of smaller models outperforms one large judge, shows less intra-model bias, and costs less. Aggregation is by union rather than majority: a concern raised by one reader survives, because in a safety-asymmetric setting the cost of a missed problem is not the cost of a false alarm. Where readers disagree sharply, the video goes to a person rather than to a shelf.
Models are scored against a hand-labelled exam under a false-safe veto: a model that passes something it should have caught is disqualified, whatever its average accuracy looks like. The published headline from that exercise is that no small model is safe alone. Every open-weight model tested was disqualified on its own. That result is why Lantern is a funnel of independent readers with a human at the disagreement points, and not a single model with a good average.
Human ratings are interleaved throughout. A parent's rulings on their own shelf are recorded and become labels, which is how the exam grows.
Lantern also reads the public comments under a video and looks for parents saying something went wrong, and it shows you what it found on the why page. A comment never blocks a video by itself.
4. The hosted alpha, in one paragraph
Each family is a separate tenant with its own database. A new family starts with hand-curated preset shelves, and every preset item is still passed through that family's own dial ceilings before it can reach a child's wall. Families that go beyond the presets bring their own API keys, so screening runs on the family's own quota and the operator holds no shared corpus. What crosses out of a family is stated on the settings page in these words: "Once a day, one thing leaves your family: yesterday's totals. How many times each video was played and finished across all families, and which sources those videos came from. The totals carry no name, no child, no time of day and nothing that points back to you, and a video's count is only ever shown once at least three different families have watched it. It is how we work out which sources are worth screening more of. Nothing else about your family leaves." The same paragraph appears in four places in the product and a build check fails if the four ever drift apart.
When a family adds its own Gemini key, what it buys is pooled the same way: "When you add your own Gemini key, the screening and the watching it pays for join the shared library that every Lantern family draws from. What joins is the video and what the reader found about it. Never your child, and never what your child watched." That sentence appears in four places in the product and the how page, and the same build check holds all five to it.
5. Limits
The evidence base is one family. Every number Lantern has produced about a real child's viewing is one child's diet under Lantern's own labels, and it is not a population claim.
The labels are machine-generated and are not clinically validated. There is no institutional review, and inter-rater reliability against a second human annotator has not been computed. Pacing figures are a model's estimate made while watching, not the output of a deterministic shot detector.
Published metadata is the ceiling, not a temporary gap. The platform's developer policies require stored data to be refreshed or deleted within thirty days and restrict derived data, so Lantern refreshes and purges on that cycle and does not download video or scrape transcripts. Some properties, notably visual pacing and loudness, are therefore judged from proxies.
The literature behind the axes is largely television and social media research carried into short-form video. Direction transfers well. Magnitudes should not be assumed to.
6. What is next
The current readers are general-purpose models asked to behave like children's media raters. The direction of the next version is a judge trained for the job: open-weight vision language models fine-tuned on these axes, able to run on a parent's own hardware, with calibration and abstention treated as behaviours of the model rather than as scripts around it. A judge trained this way can also be tuned to one child and moved as that child grows, so the ceilings that make sense at five are not the ceilings still in force at nine. This is a direction of work, not a shipping commitment, and there are no results to report yet.
7. Why any of this
The argument underneath the product is in The Primer problem: owning a machine you cannot drive is not really owning it. Parents nominally control what their children watch, but the controls belong to someone else, they are coarse, and they do not explain themselves. Dials a parent sets, on a shelf a parent can inspect, with the reasoning printed, is the same argument applied to one household screen.
8. My list
Your child can save videos to their own list. They tap "+ Add to my list" on any video, and it stays there. The list now survives closing the app, and it looks the same on a second tablet opened with the same link.
The list only ever holds videos that already passed screening for your child. If a video later comes off their shelf, it comes off the list too.
Nothing about the list leaves your family. It is kept with your family's own files, next to everything else Lantern holds for you. What leaves your family once a day is unchanged: yesterday's totals, and nothing about which videos your child saved.
What would change it: excluding a video in your review queue takes it off your child's list that same night.
References
- Alruwaily, A., et al. (2020). Child Social Media Influencers and Unhealthy Food Product Placement. Pediatrics. doi.org/10.1542/peds.2019-4057
- American Academy of Pediatrics, Council on Communications and Media (2016). Media and Young Minds. Pediatrics. doi.org/10.1542/peds.2016-2591
- American Academy of Pediatrics, Council on Communications and Media (2016). Media Use in School-Aged Children and Adolescents. Pediatrics. doi.org/10.1542/peds.2016-2592
- American Academy of Pediatrics, Council on Communications and Media (2016). Virtual Violence. Pediatrics. doi.org/10.1542/peds.2016-1298
- Bandura, A., Ross, D., and Ross, S. A. (1963). Imitation of film-mediated aggressive models. Journal of Abnormal and Social Psychology. doi.org/10.1037/h0048687
- Bond, B. J., and Calvert, S. L. (2014). A Model and Measure of US Parents' Perceptions of Young Children's Parasocial Relationships. Journal of Children and Media. doi.org/10.1080/17482798.2014.890948
- Christakis, D. A., et al. (2004). Early Television Exposure and Subsequent Attentional Problems in Children. Pediatrics. doi.org/10.1542/peds.113.4.708
- Coates, A. E., et al. (2019). Social Media Influencer Marketing and Children's Food Intake: A Randomized Trial. Pediatrics. doi.org/10.1542/peds.2018-2554
- Common Sense Media. About our ratings. commonsensemedia.org
- Coyne, S. M., et al. (2011). Profanity in Media Associated With Attitudes and Behavior Regarding Profanity Use and Aggression. Pediatrics. doi.org/10.1542/peds.2011-1062
- Coyne, S. M., et al. (2018). A meta-analysis of prosocial media on prosocial behavior, aggression, and empathic concern. Developmental Psychology. doi.org/10.1037/dev0000412
- Dal Cin, S., Stoolmiller, M., and Sargent, J. D. (2013). Exposure to Smoking in Movies and Smoking Initiation Among Black Youth. American Journal of Preventive Medicine. doi.org/10.1016/j.amepre.2012.12.008
- De Veirman, M., Hudders, L., and Nelson, M. R. (2019). What Is Influencer Marketing and How Does It Target Children? Frontiers in Psychology. doi.org/10.3389/fpsyg.2019.02685
- Funk, J. B., et al. (2004). Violence exposure in real-life, video games, television, movies, and the internet: is there desensitization? Journal of Adolescence. doi.org/10.1016/j.adolescence.2003.10.005
- Garrison, M. M., Liekweg, K., and Christakis, D. A. (2011). Media Use and Child Sleep: The Impact of Content, Timing, and Environment. Pediatrics. doi.org/10.1542/peds.2010-3304
- Garrison, M. M., and Christakis, D. A. (2012). The Impact of a Healthy Media Use Intervention on Sleep in Preschool Children. Pediatrics. doi.org/10.1542/peds.2011-3153
- Grabe, S., Ward, L. M., and Hyde, J. S. (2008). The role of the media in body image concerns among women: A meta-analysis. Psychological Bulletin. doi.org/10.1037/0033-2909.134.3.460
- Harrison, K., and Cantor, J. (1999). Tales from the Screen: Enduring Fright Reactions to Scary Media. Media Psychology. doi.org/10.1207/S1532785XMEP0102_1
- Kirkorian, H. L., et al. (2009). The Impact of Background Television on Parent-Child Interaction. Child Development. doi.org/10.1111/j.1467-8624.2009.01337.x
- Kunkel, D. (1988). Children and Host-Selling Television Commercials. Communication Research. doi.org/10.1177/009365088015001004
- Laghmiche, L., Dupire, G., and Franck, D. (2024). A New Spectrum of Self-Injuries: TikTok-Linked Lesions. Cureus. doi.org/10.7759/cureus.58226
- Lillard, A. S., and Peterson, J. (2011). The Immediate Impact of Different Types of Television on Young Children's Executive Function. Pediatrics. doi.org/10.1542/peds.2010-1919
- Lillard, A. S., et al. (2015). Further examination of the immediate impact of television on children's executive function. Developmental Psychology. doi.org/10.1037/a0039097
- Mares, M.-L., and Woodard, E. H. (2005). Positive Effects of Television on Children's Social Interactions: A Meta-Analysis. Media Psychology. doi.org/10.1207/S1532785XMEP0703_4
- Mares, M.-L., and Pan, Z. (2013). Effects of Sesame Street: A meta-analysis of children's learning in 15 countries. Journal of Applied Developmental Psychology. doi.org/10.1016/j.appdev.2013.01.001
- Papadamou, K., et al. (2020). Disturbed YouTube for Kids: Characterizing and Detecting Inappropriate Videos Targeting Young Children. Proceedings of the International AAAI Conference on Web and Social Media. doi.org/10.1609/icwsm.v14i1.7320
- Pinto, J. (2026). The Primer problem. jeffpinto.com/notes/the-primer-problem
- Radesky, J., et al. (2020). Digital Advertising to Children. Pediatrics. doi.org/10.1542/peds.2020-1681
- Strasburger, V. C., and American Academy of Pediatrics Council on Communications and Media (2010). Sexuality, Contraception, and the Media. Pediatrics. doi.org/10.1542/peds.2010-1544
- Verga, P., et al. (2024). Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models. arxiv.org/abs/2404.18796
- YouTube API Services Developer Policies. developers.google.com/youtube/terms/developer-policies
Written 2026-08-29. Updated whenever the method changes.