How I Chose Images I Could Use Without Guessing
- Dave "The Barber" Burdick

- Jul 2
- 4 min read

Before I could decide which images belonged in my first LoRA dataset, I had to answer a more basic question: which images could I use at all?
The internet is full of old illustrations. That does not make them free to use. An image may be easy to download, widely reposted, or described as “copyright free” and still have unclear ownership or restrictions. I did not want to build the project on assumptions I could not later explain.
So I made the first rule simple: every training image needs a traceable source and a clear rights statement.
What Counts as Clear Enough?

For this experiment, I am limiting the dataset to images marked Public Domain, CC0, or otherwise offered for unrestricted reuse by a credible museum, library, archive, or government institution.
CC0 is the cleanest label for this purpose. It is designed to allow copying, modification, distribution, and commercial use without asking permission. A Public Domain Mark indicates that a work has been identified as free of known copyright restrictions, although copyright status and other rights can vary by country.
That last part matters. I am not claiming that a label makes every possible use risk-free everywhere in the world. I am building a documented, rights-conscious dataset based on the best information supplied by the institutions holding the images.
It is tempting to look at a nineteenth-century drawing and assume the problem is solved. The original artwork may be old enough to be in the public domain, but I still want to know where the digital image came from and what the institution says about its use.
A random copy from Pinterest or a blog does not give me that chain of information. Even Wikimedia Commons is more useful as a discovery tool than as the final stop. Whenever possible, I want to return to the original museum or archive page and record what it says.
For every accepted image, I plan to save:
• the creator or attributed creator
• the title or identifying description
• the approximate date
• the holding institution
• the item-page URL
• the stated rights status
• the date I accessed it
• a screenshot or copy of the rights notice
• the original downloaded file
If any of that is missing or contradictory, the image goes into a review folder instead of the training set.
Free Does Not Mean Useful
I also looked at a free sample image set from Shutterstock that included food, travel, interiors, people, flowers, and graphic art. Those files may have been available for a particular approved purpose, but availability alone did not make them appropriate for this experiment.
The distinction is important: an image can be legally available and still be wrong for the dataset. Rights review answers whether I can consider the image. Style review answers whether I should use it.
The opposite is also true. I may find a perfect fantasy drawing on an auction site or personal blog, but if I cannot trace the file back to a trustworthy rights statement, it does not belong in this first training run.
Keeping the Proof With the Image
I created a source ledger so each image has its own record. The filename, source page, rights status, saved proof, and any processing notes will all use the same image ID.
That may sound excessive for a small dataset. But links disappear, websites change their wording, and files become separated from the pages where they were found. Saving the evidence now means I will not have to reconstruct the history of an image months later.
It also makes the experiment more useful as a case study. I will be able to explain not only what entered the model, but why I believed each image was eligible.

The Working Standard
My standard is not “I found it online.” It is not “the artist died a long time ago.” It is not even “someone else says it is public domain.”
The standard is: I can trace the image to a credible source, read the rights statement attached to it, preserve that evidence, and explain my decision.
That approach may eliminate images I would otherwise like to use. Good. A rights-conscious dataset should be shaped by what can be verified, not by how badly I want a particular picture.
Next, I can move to the second filter: choosing images that do not merely pass the rights test, but actually teach the specific pen-and-ink style I want the LoRA to learn.
Sources consulted
Creative Commons, CC0 1.0 Universal
Creative Commons, Public Domain Mark 1.0 Universal
The Metropolitan Museum of Art, Open Access at The Met
U.S. Copyright Office, Duration of Copyright
U.S. Copyright Office, Definitions FAQ
Pew Research Center, When Online Content Disappears
The Data Provenance Initiative, optional if you link “source ledger”



