We Made a Hands-Free TikTok in Dojo — One Sentence, Forty Seconds, No Keyboard

Val talked. Dojo generated, voiced, branded, and stitched. The first fully hands-free 40-second vertical TikTok we shipped from Dojo Workspace — no keyboard as the instrument.

By Val Neekman, with Dojo
We Made a Hands-Free TikTok in Dojo — One Sentence, Forty Seconds, No Keyboard

This is the first fully hands-free mobile video we shipped with Dojo Workspace.

Val said, out loud: make a forty-second piece for TikTok, Reels, and Shorts. Use the image, video, and audio engines. Focus. Do it.

No timeline. No keyboard as the main instrument. Talk, listen, correct, stitch, brand, export. The finished artifact is vertical — 9:16, 720×1280, 40.2 seconds, Eve on the voice, a seamless music bed, and the Dojo mark holding the first two seconds and the last three.

This is how it actually happened.

Final artifact · 9:16 · 720×1280 · 40.2s · H.264 + AAC · Eve VO + logo in/out

What we set out to prove

Dojo is not a single-model toy and it is not a prompt box with a render button glued on.

It is a workspace. You bring the LLM you already pay for. You talk. It plans. It generates. You stay in the room as director — reject a last shot, demand the logo, pick Eve’s voice, ask for a longer music bed, send it back when the proportions lie.

If that loop cannot produce a real TikTok, the product is still a demo.

On 24 August 2026 it produced one.

What you need

Nothing exotic. This is the stack we actually used:

  • Dojo Workspace running locally, voice on, Advanced Media Playbook available.
  • A video model wired in the lane (we used the image-to-video / text-to-video path for seven short cinematic beats).
  • Cloud voice — Val enabled premium voices and named Eve for the VO.
  • A brand card — the Dojo mark plus heydojo.ai, held two seconds at the open and three at the close.
  • Taste. The machine will generate. A human still says “the last segment is wrong” and means it.

You do not need an editor, a camera, or a second person in the chair. You need a sentence and the willingness to stay until the last frame is honest.

16:9 skyline still from the first generated flyover segment
Segment 1 — skyline reveal. This is the first picture the model gave us.

The conversation, not the timeline

The session title in Dojo is still the first line Val spoke:

Hey dojo what I want is I definitely want to start 40 some second video for TikTok…

Session id 0955eb59-…. Primary lane. Project: Mobile.

Dojo asked one preference — vibe — and Val picked cinematic city flyover at dusk. After that the work was spoken direction:

  1. Lock the format. Vertical 9:16. Target ~40 seconds. Clip generators arrive in short bursts, so we planned seven 5-second segments and a stitch.
  2. Generate the spine. Skyline → canyon dive → highway light trails → tower crane → rooftop gardens → tower-gap climax → pullback finale.
  3. Voice. “Use Eve’s voice. I enabled Cloud Voices.” Premium TTS, not a local fallback.
  4. Brand the gate. Val rejected a cut that did not end on the mark. Logo on black. URL at the bottom. Hold two or three seconds. Then the same card at the start for two seconds.
  5. Sound. Keep each segment’s own atmosphere through the stitch. Then pull the longest, cleanest bed from one scene and let it carry the whole piece so the music does not die in the middle.
  6. Fix the lie. The last segment’s proportions were wrong. Val sent a screenshot. We regenerated that beat until the geometry held.
  7. Export. 720p for this first ship (faster next time can be 480p drafts). H.264 + AAC. File lands at media/exports/final-artifact-tiktok-reels/2026-08-24-neon-city-final-VERTICAL-40s.mp4.

That is the whole “edit.” Spoken notes. Regenerations. A stitch. A brand hold.

16:9 highway light-trail still from the flyover
Segment 3 — highway light trails. The bed we looped came from a long, clean scene like this.

The seven beats

BeatSecondsPicture
10–5Skyline reveal at dusk
25–10Canyon dive between glass
310–15Highway light trails
415–20Skyscraper and crane
520–25Rooftop gardens
625–30Tower-gap climax
730–35Epic pullback
Brandfirst 2s + last 3sDojo mark + heydojo.ai

The generators do not give you a 40-second master in one call. They give you honest short shots. The craft is deciding the order, keeping identity across cuts, and refusing the beat that breaks scale.

Splashy 16:9 neon downtown canyon
The canyon energy — wet street, stacked signs, the dive between towers.

Voice, logo, music — the three things amateurs skip

Voice. Eve reads a short VO. Not a lecture. A presence. Hands-free only works if the audio track is a person, not a subtitle file you forgot to burn in.

Logo. Val was blunt: end on the mark and the download. Two seconds in, three seconds out, black field, URL low enough to survive TikTok’s UI chrome (keep faces and type out of the bottom ~15%).

Music. First stitch let the beds die at each cut. Val asked us to lift the longest clean atmosphere and repeat it under the whole film so the city feels like one place.

Those three notes are why this is a post, not a render dump.

16:9 finale still — epic pullback over the neon city
Segment 7 — the pullback. The last live picture before the mark.

Specs we actually shipped

  • Duration: 40.23s
  • Frame: 720×1280, 9:16, 30 fps
  • Codecs: H.264 (avc1) + AAC (mp4a), ~20 MB
  • Voice: Eve, premium cloud TTS
  • Brand: Dojo card in / out, heydojo.ai
  • Safe zone: type and mark held out of the lower UI band
  • Captions we wrote the same afternoon:
The cities of the future are already glowing. One sentence became this — imagined, voiced & built with Dojo. Download at heydojo.ai

Hashtags, brand first, then three topic tags:
#HeyDojo #DojoWorkspace #DojoSolo #DojoDuo #Dojo #aivideo #neoncity #fyp

Why this matters

Anyone can paste a prompt into a website and get a clip.

Almost nobody can sit in a chair, talk, reject the bad ending, name the voice, demand the mark, and leave with a file that is legal for three platforms the same day — while the same workspace also writes code, reviews a branch, and keeps a dated content system (content/2026/08/2026-08-24.md).

That is the product.

Bring your model. Grok, Claude, Gemini, OpenAI, Alibaba, local — Dojo is the hands and the face. The brain is the partner you already chose. Canada first. The world next. Charity starts at home.

If you want to try the same loop: heydojo.ai.

Press play on the video above. That is not a concept. That is Tuesday.

https://heydojo.ai/r/hands-free(short URL)