Turn your document
into a training set
Upload a PDF, DOCX or text file. It gets split into passages, each passage is run through 7 different question-and-answer patterns, and you get a dataset ready for fine-tuning. Answers use facts from your document only. Some questions are invented on purpose, so the model also learns to correct a wrong belief and to decline an off-topic request.
Early access: accounts are approved one by one. Sign up and we will add credits to yours and let you know. No card required.
How it works
Three screens, twelve steps behind them. You only touch the first and the last.
Upload a document, state your goal
Say in a paragraph what the model is for, who will use it, and what language it answers in. The model reads your whole document and writes its own brief.
Split into 800–1000 token passages
Passages break between paragraphs whenever possible; a paragraph or table longer than one passage is split at the nearest line or sentence. Each passage goes through 7 separate prompts and yields up to 30 examples.
Download as one file
Examples are merged in passage and pattern order. Chat JSONL, ShareGPT, Vertex AI or Alpaca: whichever your trainer reads. You can download what exists while the run is still going.
7 patterns for every passage
One kind of question teaches a model to recite. The same passage is worked from 7 angles, so the dataset teaches behaviour as well as facts.
knowledgeknowledgereasoningreasoningmulti_turnmulti-turncalibrationuncertaintynegativecounter-examplestylestylerefusalrefusalWhat the output looks like
Every example is a conversation: a user question and an assistant answer. Below is an illustrative example for a leave-policy document.
How you read it
What happens to annual leave after five years of service?
An employee who completes five years moves from 14 to 20 days of annual leave. The increase applies in the leave period following the year the tenure is completed.
How it downloads
Chat JSONL (OpenAI messages format). ShareGPT, Vertex AI and Alpaca are one click away.
{"messages":[
{"role":"system",
"content":"You are an HR assistant..."},
{"role":"user",
"content":"What happens after five..."},
{"role":"assistant",
"content":"...moves from 14 to 20 days."}
]}What a typical run looks like
Estimated from our production runs with Gemini 3.5 Flash-Lite, for a 60,000-token document (roughly 150 pages). Runs take longer when several share the queue.
- Passages
- 67
- Examples
- ~1,930
- Time
- ~12 min
- Credits
- 469
1 credit = 1 model call. Credits are estimated from document length, about passages × 7, and deducted when the run starts.
What we don't do
We do not score the examples for quality. Only the shape is checked: a conversation starts with the user, alternates, and ends with the assistant; malformed examples are dropped. No judge model, no filtering pass, no deduplication.
That is a deliberate choice, not an omission. Every filtering layer we tried multiplied cost and time without a measurable gain. The prompts carry the quality — and you can download the whole output and read it yourself.
Request access
Request access, and once your account is approved you can upload a document, download the dataset, and judge it yourself.
Request early access