We Made a Hands-Free TikTok in Dojo — One Sentence, Forty Seconds, No Keyboard
Val talked. Dojo generated, voiced, branded, and stitched. The first fully hands-free 40-second vertical TikTok we shipped from Dojo Workspace — no keyboard as the instrument.

This is the first fully hands-free mobile video we shipped with Dojo Workspace.
Val said, out loud: make a forty-second piece for TikTok, Reels, and Shorts. Use the image, video, and audio engines. Focus. Do it.
No timeline. No keyboard as the main instrument. Talk, listen, correct, stitch, brand, export. The finished artifact is vertical — 9:16, 720×1280, 40.2 seconds, Eve on the voice, a seamless music bed, and the Dojo mark holding the first two seconds and the last three.
This is how it actually happened.
What we set out to prove
Dojo is not a single-model toy and it is not a prompt box with a render button glued on.
It is a workspace. You bring the LLM you already pay for. You talk. It plans. It generates. You stay in the room as director — reject a last shot, demand the logo, pick Eve’s voice, ask for a longer music bed, send it back when the proportions lie.
If that loop cannot produce a real TikTok, the product is still a demo.
On 24 August 2026 it produced one.
What you need
Nothing exotic. This is the stack we actually used:
- Dojo Workspace running locally, voice on, Advanced Media Playbook available.
- A video model wired in the lane (we used the image-to-video / text-to-video path for seven short cinematic beats).
- Cloud voice — Val enabled premium voices and named Eve for the VO.
- A brand card — the Dojo mark plus
heydojo.ai, held two seconds at the open and three at the close. - Taste. The machine will generate. A human still says “the last segment is wrong” and means it.
You do not need an editor, a camera, or a second person in the chair. You need a sentence and the willingness to stay until the last frame is honest.

The conversation, not the timeline
The session title in Dojo is still the first line Val spoke:
Hey dojo what I want is I definitely want to start 40 some second video for TikTok…
Session id 0955eb59-…. Primary lane. Project: Mobile.
Dojo asked one preference — vibe — and Val picked cinematic city flyover at dusk. After that the work was spoken direction:
- Lock the format. Vertical 9:16. Target ~40 seconds. Clip generators arrive in short bursts, so we planned seven 5-second segments and a stitch.
- Generate the spine. Skyline → canyon dive → highway light trails → tower crane → rooftop gardens → tower-gap climax → pullback finale.
- Voice. “Use Eve’s voice. I enabled Cloud Voices.” Premium TTS, not a local fallback.
- Brand the gate. Val rejected a cut that did not end on the mark. Logo on black. URL at the bottom. Hold two or three seconds. Then the same card at the start for two seconds.
- Sound. Keep each segment’s own atmosphere through the stitch. Then pull the longest, cleanest bed from one scene and let it carry the whole piece so the music does not die in the middle.
- Fix the lie. The last segment’s proportions were wrong. Val sent a screenshot. We regenerated that beat until the geometry held.
- Export. 720p for this first ship (faster next time can be 480p drafts). H.264 + AAC. File lands at
media/exports/final-artifact-tiktok-reels/2026-08-24-neon-city-final-VERTICAL-40s.mp4.
That is the whole “edit.” Spoken notes. Regenerations. A stitch. A brand hold.

The seven beats
| Beat | Seconds | Picture |
|---|---|---|
| 1 | 0–5 | Skyline reveal at dusk |
| 2 | 5–10 | Canyon dive between glass |
| 3 | 10–15 | Highway light trails |
| 4 | 15–20 | Skyscraper and crane |
| 5 | 20–25 | Rooftop gardens |
| 6 | 25–30 | Tower-gap climax |
| 7 | 30–35 | Epic pullback |
| Brand | first 2s + last 3s | Dojo mark + heydojo.ai |
The generators do not give you a 40-second master in one call. They give you honest short shots. The craft is deciding the order, keeping identity across cuts, and refusing the beat that breaks scale.

Voice, logo, music — the three things amateurs skip
Voice. Eve reads a short VO. Not a lecture. A presence. Hands-free only works if the audio track is a person, not a subtitle file you forgot to burn in.
Logo. Val was blunt: end on the mark and the download. Two seconds in, three seconds out, black field, URL low enough to survive TikTok’s UI chrome (keep faces and type out of the bottom ~15%).
Music. First stitch let the beds die at each cut. Val asked us to lift the longest clean atmosphere and repeat it under the whole film so the city feels like one place.
Those three notes are why this is a post, not a render dump.

Specs we actually shipped
- Duration: 40.23s
- Frame: 720×1280, 9:16, 30 fps
- Codecs: H.264 (
avc1) + AAC (mp4a), ~20 MB - Voice: Eve, premium cloud TTS
- Brand: Dojo card in / out,
heydojo.ai - Safe zone: type and mark held out of the lower UI band
- Captions we wrote the same afternoon:
The cities of the future are already glowing. One sentence became this — imagined, voiced & built with Dojo. Download at heydojo.ai
Hashtags, brand first, then three topic tags:#HeyDojo #DojoWorkspace #DojoSolo #DojoDuo #Dojo #aivideo #neoncity #fyp
Why this matters
Anyone can paste a prompt into a website and get a clip.
Almost nobody can sit in a chair, talk, reject the bad ending, name the voice, demand the mark, and leave with a file that is legal for three platforms the same day — while the same workspace also writes code, reviews a branch, and keeps a dated content system (content/2026/08/2026-08-24.md).
That is the product.
Bring your model. Grok, Claude, Gemini, OpenAI, Alibaba, local — Dojo is the hands and the face. The brain is the partner you already chose. Canada first. The world next. Charity starts at home.
If you want to try the same loop: heydojo.ai.
Press play on the video above. That is not a concept. That is Tuesday.