Revesery
Dashboard
Home
Bookmarks
Messages
Tools
Explore
Social Media
VideoBulk VideoProfile PictureSlide ShowSound / AudioUnfollowersDouyin
VideoBulk Video
VideoStoriesBulk VideoProfile PictureSlide ShowUnfollowersStalker CheckSoon
MP4 · VideoMP3 · Audio
VideoUnfollowers
WhatsApp
More tools
PinterestThreadsTeraboxVidey
Utilities
Case Converter
Cookie Converter
Review Calculator
Deep Voice Checker
Fake SNBT
Watermark KTP
Surat Izin
Add bookmarks
Revesery
⌘K

Most read

Nothing published yet — type a topic and we’ll dig.
ShareLogin
Back to feed
AnonymousAnonymous
@anonymous8771 · Aug 14, 2026 · 6 views
#AI#News

OpenWALDO Brings Open Source Transparency to AI Training Data

Ever wished you could trace exactly where an AI model's training data came from? The creator of CentOS and Rocky Linux is building a dataset that makes that possible.

OpenWALDO packs about 202.8 billion tokens from government records, academic research, mailing lists, and public domain books — and every token can be traced back to its source.

OpenWALDO Brings Open Source Transparency to AI Training Data

What Happened

Gregory Kurtzer, the founder of CentOS and Rocky Linux, just launched OpenWALDO. It is an open, transparent dataset for AI training, built on roughly 202.8 billion tokens from sources like government records, academic research, mailing lists, and public domain books.

The core idea is simple: every token is traceable to its source, every source is documented, and every license is accounted for. No more wondering whether a rare book got scraped without permission.

Kurtzer argues Linux won not because it was certified safe, but because people could inspect it, fork it, and have the whole community review and test it. He thinks the same logic should apply to AI. Any lab or company can take the OpenWALDO corpus, verify the sources, add its own proprietary data, and build a model with a clear, auditable lineage.

The project is funded by CIQ, Kurtzer's company behind Rocky Linux, and it is looking for contributors.

Why It Matters

Most models today are built on datasets assembled from random web scraping, with sources that are hard to pin down. OpenWALDO offers an alternative: a base corpus where the provenance of every token is known. If labs and companies build on it, we get AI models with real accountability — the same way open source transformed operating systems. The open question is whether that approach can win in AI the way it did with Linux.

Liked this share? Revesery is where people swap what they're actually building.

Join with Google
Be the first to sayBe first

Does this still work?

Nobody's checked yet

Sign in to tell everyone how it went.

Continue with Google

Be the first — one tap saves the next person an hour.

It takes 3 reports in 30 days to set the status.

Comments

Join the conversation — sign in to comment.

Sign In Now

No comments yet — start the conversation!

More shares you might like

AIHow to Get Free $200 Funds on Digital OceanAlex Ruiez · 1y · 692 viewsToolsHow to Deploy to Your VPS With MidrepoAnonymous · 2h · 7 viewsHow I Got 600 Followers in 2 DaysRodney · 1y · 443 viewsToolsHow to Get 1 Year of Atomesus Pro for FreeAnonymous · 1d · 22 viewsToolsHow to Check If a New GrabFood Account Is FlaggedAnonymous · 2d · 23 viewsToolsHow to Make AI Gold Extraction Videos in Google FlowAnonymous · 5d · 51 views

Site footer

Revesery

Empowering people to share valuable insights, discover hidden information, and connect with an amazing community of learners and experts.

  • 201Members
  • 159Shares published

Explore

  • Explore shares
  • Trending now
  • Tools
  • Top contributors

Company

  • About us
  • Contact
  • Advertise with us
  • System status
© 2026 Revesery
  • Privacy
  • Terms
  • Trust & Safety
HomeExploreShareToolsProfile