Recording & Auto-Writing (Mining)
Record a quick voice memo, and VoiceDrop automatically transcribes it and "mines" it into one or more publish-ready articles.
This is the core VoiceDrop workflow: no typing, no organizing. Just say what's on your mind into your phone, and the app handles everything else — transcription, structuring, and turning it into a finished article.
On the "My Recordings" home screen, a red record button floats at the bottom. Tap it to enter the full-screen recording view and start recording. When you're done, tap Stop, and the recording is automatically queued for upload. Once uploaded, the server first transcribes your speech into text (Transcribing), then has AI mine your spoken words into a fully structured article (Mining) — with a title, an opening, and an ending, reading like an article you wrote yourself.
You don't have to watch any of this happen. Each recording moves through the statuses Uploading → Queued → Transcribing / Mining → Done in the list. Once done, tap it to hear the original audio, read the article, and publish with one tap. If no speech is detected in the recording, it's marked No Speech; if it's too fragmented to form a coherent piece, it shows "No article yet" — try again with a more complete thought.
How to use it
- Open the app and stay on the "My Recordings" tab at the bottom.
- Tap the red record button at the bottom. The screen switches to a full-screen recording view and recording starts immediately — "Recording" at the top, a large timer and a bouncing waveform in the middle. (Note: recording is tap-to-start, not press-and-hold. The "hold to speak" label on the red button is a different feature — holding the red button lets you give voice commands to the whole list, like "delete the second one".)
- To finish, tap the square Stop button in the middle. The recording view closes, you're back at the list, and the recording shows "Uploading".
- Want photos with your recording? Tap the light-colored camera icon on the right side of the recording view to snap a few shots. The photos upload together with the recording and get woven into the matching text by AI (taking photos doesn't interrupt recording).
- Then do nothing. The recording moves automatically from "Queued" through "Transcribing" and "Mining", and finally turns into a green "Done".
- Tap a "Done" recording to play back the original audio on top and read the mined article below. The ⋯ menu in the top right lets you "Publish to WeChat Drafts" or share.
- To edit the article: open the article view, hold the "Hold to Speak · Edit" talk bar at the bottom and speak your edit (like "delete the second paragraph"), then release — done.
- To have AI mine it again from scratch: in the list, long-press the "Done" badge on a finished recording and choose "Rewrite". It reuses the existing transcript and re-mines with the original logic (the output may differ from before).
Details & rules
- How to start recording: tap the red record button on the home screen; the full-screen view starts recording automatically. Tap "Stop" to finish. Actual recording is tap-based, not press-and-hold.
- How many articles per recording: the default is "merge aggressively, fewer but richer" — one voice memo usually produces just 1 article. Only when you clearly jump between unrelated topics does it split into 2–3 articles, each able to stand on its own. The cap is 3.
- Only facts you actually said: the AI has a hard rule — it only uses content from the transcript, never invents or embellishes. Length follows the content: a few sentences can become a short piece; it never pads for word count.
- Follow-up questions: after each article is written, the AI picks out the "thinnest" spots that only you would know and asks 1–3 short follow-up questions (real numbers, real names, judgments left unexplained). The questions are stored separately and never enter the article body; you can answer by voice and let it fill things in.
- Status badges: Uploading (red) → Queued (amber) → Transcribing / Mining (spinner) → Done (green). No detected speech shows "No Speech"; blocked recordings show "Out of Credits" or "Recording Too Long".
- When articles get written: the server processes automatically once an hour (and also triggers on upload). Transcribing a long recording may take several passes, so just check back a bit later.
- Offline & retry: you can record without a connection. Recordings are saved to a local queue on your phone and upload automatically once you're back online. A failed upload retries automatically (up to 3 times, at 1.5s and 3s intervals); failed files stay safely on your phone and get retried later. The moment the network comes back, uploads resume automatically.
- Recording length limit: a single recording can be up to 3 hours. Beyond that it's marked "Recording Too Long" and not processed.
- Transcript too short: if the transcript is under 20 characters, it's marked "Too Short" and no article is written.
- Credits (cost): new users get a one-time gift of 200 credits (valid for 1 year); 23 credits = ¥1. Transcription is billed by duration (about ¥0.8/hour); mining is billed by AI usage. When your balance hits 0 (out of credits), new recordings stop being processed until you top up.
- Appending / editing logic: each recording maps to its own article(s) — there's no "append a new recording to an old article". To change an existing article, use the hold-to-speak edit bar; to redo it entirely, use "Rewrite" to mine again.
- Recording from a tag page: if you start a recording from inside a tag page, the mined article gets that tag by default.
FAQ
How many articles can one recording produce?
Why does my recording still say "Queued", and how long until it's done?
Can I record without an internet connection?
Why does my recording show "No Speech" or "No article yet"?
Follow-up Questions
After an article is mined, the AI asks about its weakest spots — one or two details only you know. Hold to speak your answer, and the article grows richer with every reply.
You record a quick memo, and VoiceDrop mines it into a fully structured article. But recordings often leave out key information "only you know" — a real number, a person's name, a concrete scene, or a judgment you touched on but never unpacked. Fill those in, and the article truly stands up.
That's what follow-up questions are for. After each article takes shape, the AI finds its thinnest spots and asks 1 to 3 short, specific questions — like an editor pressing the author for details, not small talk. No typing needed: hold to speak your answer, and the information gets woven into the most relevant paragraphs automatically.
Follow-ups are a lightweight feature that's on by default. It stays out of the way — normally tucked away, showing only as a star with a number badge next to the talk bar on the recording detail page, reminding you how many questions are unanswered. Open it when you feel like answering; it never nags. Once you've answered or skipped them all, the star disappears.
How to use it
- Record as usual, wait for VoiceDrop to mine it into an article, and open the article's detail page.
- If the AI has questions for you, a star button lights up on the right of the "Hold to Speak · Edit" bar at the bottom. The small number badge in its corner is the count of unanswered questions.
- Tap the star, and the follow-up card expands from the bottom, wrapping around the talk bar: above it, "Question N/M" (current / total) and the question itself; below, a segmented progress bar.
- To answer, hold the talk bar (its label changes to "Hold to Speak · Answer"), speak your answer, and release to send. You don't need to echo the question — just state the facts you know.
- After you release, the AI weaves your answer into the most relevant paragraphs. The updated paragraphs highlight in yellow for a few seconds so you can see where the article grew, then it automatically flips to the next question.
- Don't want to answer one? Tap "Skip" in the card's top right corner to move on.
- When everything is answered or skipped, the card and star tuck away automatically, back to the normal talk bar.
- Want more questions? Just hold and say "ask me a few more" or "anything else you want to know" — it generates new questions on the spot and the star lights up again. (Note: while the follow-up card is open, anything you speak is treated as an answer. To ask for more questions, collapse the card first.)
- Don't want this feature at all? Turn off the "Follow-up after writing" switch in Settings (subtitle: AI asks a detail or two to thicken the article).
Details & rules
- The AI asks 1 to 3 questions per article, all about specific information that can't be inferred from the recording — things only the author knows (real numbers, real names, real scenes, unexplained judgments). If nothing is worth asking, it doesn't ask.
- Follow-up questions never enter the article body or version history. Publishing to WeChat, generating a share page, posting to the community, exporting to Xiaohongshu — every outlet is naturally free of these questions. Only you see them, inside the app.
- Answers travel through exactly the same channel as regular voice edits — each answer is just an ordinary edit command, with no special waiting screen. Once sent, you get the familiar queue bubble and "editing" flow.
- Per-question progress: green = answered, orange = current, gray = unanswered or skipped.
- Unanswered questions expire automatically after 7 days — they don't hang around forever.
- Follow-ups are per-article: each article has its own set of questions, and the star badge shows the unanswered count for the article you're viewing.
- The yellow highlight after an answer just shows "what changed where" — it fades in a few seconds and doesn't block anything.
FAQ
Why do some articles get follow-up questions and others don't?
Do I have to answer each question word by word?
Will the follow-up questions get published with my article, where others can see them?
How do I get the AI to ask me more questions?
Editing Articles by Voice
Open a mined article, hold the talk bar at the bottom, and say what to change — the AI edits accordingly: body text, title, deleting paragraphs, swapping images, all of it.
A mined article isn't the end of the road. On any article's reading page, there's a persistent WeChat-style talk bar at the bottom — normally labeled Hold to Speak · Edit. Hold it, say what to change and how, release — and the AI understands your words and edits the article directly. No typing, no hunting for a cursor in a text box.
It edits the article you're currently reading. So you can point precisely, every line of the body is numbered "Line N / Image M" at the start — say "delete line 3", "change line 5 to…", "change the title", and the AI locates targets strictly by the numbers on screen, never miscounting. While you speak, the transcript shows live at the top of the screen, with locator phrases like "Line N / Image M" highlighted so you can confirm it heard the right line or image.
You can issue edits one after another. Each command joins a queue and runs in order — the one being executed shows a bouncing pencil, and the rest wait behind it. When an edit lands, the changed lines flash with a highlight and slowly fade, so you can see at a glance what changed where. Every voice edit is automatically saved as a new version, so if something goes wrong you can always undo.
How to use it
- Open a mined article's reading page. A persistent talk bar sits at the bottom, labeled Hold to Speak · Edit.
- Hold the talk bar and speak your edit, e.g. "delete line 2", "rewrite line 5 as…", "change the title to…", "add a paragraph after line 3 about…". Keep your finger down while speaking.
- As you speak, watch the dark bubble at the top of the screen — it shows the live transcript, with your "Line N / Image M" locators highlighted so you can confirm the target.
- When you're done, release to send (to abandon before releasing, swipe up to cancel as prompted). The command joins the queue and starts executing.
- Want several changes in a row? Don't wait — hold and speak the next one. New commands stack in the queue and run one by one. The one executing shows a bouncing pencil icon; queued ones show a little clock.
- When an edit finishes, the AI replies above the talk bar (on success, usually "Done"), and the changed lines flash briefly with a highlight.
- For more precise edits, you can long-press a paragraph — a small menu pops up with "Rewrite this paragraph", "Insert image", and more (plus a local "Copy"). Long-pressing a generated illustration lets you change that image's style. Your choice is handed to the same voice-edit queue.
- Not happy with an edit? Use the Undo / Redo buttons in the toolbar to step back or forward (they appear once there's more than one version).
Details & rules
- It always edits the article you're looking at. The "Line N / Image M" numbers in the body exist only to align you and the AI on targets — they're never written into the article itself.
- Commands you can give: delete a line, change a line, insert a paragraph after a line, change the title; you can also have it rewrite the whole piece. For images: swap an image or generate a new one. Deleting a paragraph means deleting the corresponding line N.
- Merging two articles or deleting a whole article — cross-article operations — don't happen on the single-article page. Do them from the "My Recordings" list by holding the red button and speaking with article numbers (e.g. "merge article two and article three"). Merging saves a new article and keeps the originals; deleting asks you to confirm first.
- Queued, serial execution: multiple commands don't run at once — they enter one queue and execute in order. The active one is highlighted; the rest wait. The server holds this queue and is the true source of authority.
- Versions are kept: every voice edit automatically writes a new version (source tagged as agent), with a maximum of 10 versions kept — the oldest gets dropped beyond that. Undo / Redo only moves the "current version" pointer without writing a new version; if you undo and then make new edits, the undone "future" versions are discarded, just like git.
- Feedback after edits: the AI leaves one line of reply above the talk bar — a bright dot icon on success, a red warning icon on error. The reply doesn't auto-dismiss until replaced by the next one or you tap elsewhere. Changed lines glow for a few seconds and fade.
- Nothing is lost on disconnect / backgrounding / force-quit: commands you've spoken but the server hasn't confirmed are stored locally and resume automatically on reconnect. Stable command IDs deduplicate the resume, so the same edit never applies twice. Text-only commands survive even a force-quit; in the rare case an image edit is interrupted mid-flight, you'll need to say it again (image edits don't silently resume).
- When a command fails: the reply turns into a red warning with a hint, and that command is removed from the queue — just rephrase more clearly and say it again. If the network dropped mid-way, the command stays in the queue and is sent again after auto-reconnect; no manual resend needed.
FAQ
Can I undo a bad edit?
Can I dictate several edits in one go?
How do I target an exact line or image?
If the network drops mid-edit, or I quit the app, do I lose my changes?
Images
Take photos or import from your library inside VoiceDrop — let the AI write from what it sees, or generate cover images, cartoon explainers, and illustrations in various styles for existing articles.
VoiceDrop isn't just about voice — photos are material for the AI too. You can snap photos while recording to attach what's in front of you, or add photos to any finished article later, taken on the spot or picked from your library.
If a recording contains only photos and no speech, VoiceDrop doesn't discard it as "No Speech" — it automatically switches to photo mode: the AI observes the photos and writes a minimal photo essay for you. In other words, a single photo can become a short article on its own.
Photos can also flow the other way — you can have the AI draw for your article: long-press a paragraph to generate a WeChat Official Account cover image, or a cartoon explainer that helps readers grasp the article's structure at a glance. Long-press an existing illustration to redraw it in a different style — cartoon, watercolor, sketch, oil painting, film, or advertising. All drawing jobs join the same queue as voice edits; the AI processes them one by one, and the article updates automatically.
How to use it
Shoot while recording (photos during a recording)
- On the recording screen, tap the faint camera icon on the right to open the square camera.
- The viewfinder is square, with a rule-of-thirds grid. Tap the white shutter button to shoot; photos line up in a filmstrip at the bottom.
- The library icon in the bottom left picks photos from your system library (up to 9 at a time); the bottom right switches between front and rear cameras.
- To remove a photo, tap the ✕ in its top right corner.
- Tap "Done" in the top right when finished — the photos attach to the recording and upload with it.
Photos only, no speech — straight to an article
- Take or pick photos as above; you don't have to say a word.
- When the recording ends, the AI sees photos but no speech and automatically enters photo mode, observing the photos and writing a minimal photo essay for you.
Insert photos into an existing article
- Open an article and tap the "Insert photos" icon (film-strip style) in the top toolbar.
- The same square camera opens — shoot on the spot or pick from your library.
- Tap "Done"; the photos upload and the AI places each one near the paragraph it fits best.
Have the AI draw for your article (long-press a paragraph)
- In edit mode, long-press any paragraph to open the action menu.
- Choose "Insert image" — it has two options:
- WeChat cover image: a 2.45:1 banner cover placed at the very top of the article, with a title distilled from the article (about 6–10 characters).
- Cartoon explainer: a flat cartoon-style diagram inserted where it best aids understanding, making the article's structure clear at a glance.
- The menu also has "Rewrite this paragraph" (more concise / more casual / more formal / expand a bit) and "Copy".
Restyle an existing illustration (long-press the image)
- Long-press an AI-generated illustration in the article to open the "Image style" menu.
- Options: cartoon, advertising, watercolor, sketch, oil painting, film.
- Pick one, and the AI redraws the image in that style while keeping the composition and subject intact.
Details & rules
- Library picker limit: up to 9 photos per pick.
- Photo format: every photo is center-cropped to a square and scaled so the longest edge is at most 1080 pixels, saved as JPEG (under 900KB each). So illustrations in the article match exactly what you saw in the viewfinder — always 1:1.
- Camera shots are also saved to your phone's photo library: the saved copy is the same square version as the viewfinder. Photos picked from the library are already there and aren't duplicated. (If library write permission is denied, they simply aren't saved — everything else still works.)
- Where photos live: they upload to cloud storage under your own account, at paths like
photos/<session-timestamp>/<second>-<random-suffix>.jpg. Each photo carries its own marker and is placed in the article by that marker, not by sequence number. - Photo-mode limits: articles generated from photos alone are minimal photo essays, and photo mode does not generate follow-up questions (there's no dictation to answer about).
- Drawing shares one queue with voice edits: cover images, cartoon explainers, and restyle commands from the long-press menu all join the same edit queue. The AI processes them serially, showing "editing" while it works; the article updates automatically when done, with undo / redo available.
- Cover vs. explainer conventions: WeChat cover images are fixed at a 2.45:1 banner; a cartoon explainer's aspect ratio follows the content (landscape, portrait, or square), aiming for "understood at a glance".
- These menu items (rewrite, insert image, image style) are configured server-side and may change between releases.
FAQ
I only took photos and didn't say a word — can it still produce an article?
How many photos can I pick from the library at once?
Do photos taken with the in-app camera get saved to my phone's photo library?
What's the difference between a "WeChat cover image" and a "cartoon explainer"?
Writing Style
Teach VoiceDrop how you write, so every recording gets mined into an article that reads like you wrote it.
Your "writing style" is a style fingerprint VoiceDrop keeps for you — it records how you like to write: sentence rhythm and length, favorite words, how tight or loose your tone is, how you open and how you land — not what you write about or what you believe.
It has exactly one job, but a crucial one: every time your voice or photos get "mined" into an article, VoiceDrop folds this style into the AI's instructions, so the output tastes like you — not generic AI-speak.
Styles are stored versioned — every save adds a version, and you can switch or roll back anytime, or even mine the same recording with different versions to compare. There are three ways to get one: write or paste one yourself in Settings; use "Style Learning" to have the server distill one from articles you admire; or have Claude distill one from your published work. Once a style is in place, it applies to all future recordings — and older articles can be re-mined with any style version too.
How to use it
1. View and hand-write your own style
- Go to Settings and tap "Writing Style" in the first card (subtitle: "mimic this voice when writing").
- A full-page editor opens. If you have no style yet, it shows the hint: "No writing style yet. Paste a distilled style here, and mining will apply it to make articles sound more like you."
- Type or paste your style description into the box and tap "Save" in the top right. Every save adds a new version.
- A version bar at the top shows the current version number, style name, character count, and total versions. Tap it to expand the version dropdown and switch to any past version. Switching to an old version and saving without edits equals a rollback (no new version); editing the text and saving creates a new version.
- Tap "Cancel" in the top left to leave without saving.
2. Style Learning — distill a style from other people's articles
- From any app (a Safari page, a text selection, a PDF/Word/RTF document), use the system Share button and choose VoiceDrop.
- The "Style Dataset" panel appears: the top shows "N items collected · ~X characters", with the corpus so far and what this share added below. Web pages get their body text parsed automatically (parsing → collected; on failure it shows "link only", with a "Retry" button).
- Tap "Keep collecting" to close the panel and share in a few more pieces.
- When you have enough, tap the orange "Extract Writing Style". There's a "clear dataset after extraction" checkbox at the bottom (checked by default, so next time starts fresh).
- VoiceDrop distills in the background and jumps to "My Recordings" in the app. When done, an introduction article titled "Your writing style · name" appears.
3. Re-mine an old article in a different style (restyle)
- Open any finished article (reading page).
- Next to the date under the title, there's a style tag with a pencil icon (showing the current style version, e.g. "v8 style"). Tap it.
- The "Rewrite in a different style" panel opens: "Pick a style version and rewrite this article. The original stays; switch back anytime."
- Pick a style version and tap "Rewrite with vN" at the bottom. If this article has used that version before, it switches back instantly (free); otherwise it re-mines with that style and produces a new version.
- During the rewrite it shows "Rewriting in the new style…". The original dictation is untouched, and you can undo or switch versions anytime.
Details & rules
- Style captures "how you write", not "how you think": the distiller extracts 9 dimensions — sentence rhythm, paragraph length, vocabulary, tone, argument structure, metaphors, emotional intensity, openings and endings, and "things you never do" — anchoring each with a few real quoted sentences, and names the style in 5 characters or fewer (shown as the first line of the version).
- Style Learning has a word-count threshold: the corpus needs at least 300 characters of usable body text before extraction is allowed. Below that, the button is disabled with the hint "Not enough material yet (X characters total) — share a few articles with body text and collect 300+ characters first". Shares that carry only a title or link with no body (like book-title-only shares from reading apps) don't count.
- Too few samples get shaky: with fewer than 3 pieces in the corpus, the distilled result is flagged "the fingerprint may be unstable" — it's easy to mistake one article's quirks for your signature.
- Corpus limits: each sample keeps up to its first 4,000 characters, and one distillation feeds the model at most about 48,000 characters total — anything beyond is truncated in collection order.
- Extraction runs on the mining pipeline: tapping "Extract Writing Style" actually uploads a silent placeholder recording that triggers the server-side miner — so you can watch progress in "My Recordings" like any recording, and retry on failure. More reliable than instant extraction.
- It won't wreck anything when material is short: if the server finds the corpus under 300 characters, it does not touch your active style. Instead it gives you an explainer article — "Not enough samples, style unchanged" — listing what it received and how to add more. Your material stays in the dataset.
- Versioning and rollback: styles are stored versioned just like articles, keeping roughly 10 versions of history; switch or roll back anytime from the version dropdown.
- Multi-style comparison: the version dropdown has a "multi-style comparison" toggle, letting you check 2–3 versions (3 max). The idea: mining generates one article per style so you can flip between them at the top of the reading page.
- Re-mining costs credits: restyling an old article runs a full mining pass and costs credits normally. If the article has already used that version, switching back is free — nothing is regenerated.
FAQ
I changed my style — will my previously mined articles change with it?
How much material does Style Learning need before I can extract?
Will extracting a new style overwrite my old one for good?
A shared web page says "parse failed · link only" — what do I do?
VD Community (Share, Browse, Respond, Tip & Report)
Share your AI-mined articles with fellow creators — read each other's work, record responses, toss coins in encouragement, and withdraw or report content anytime.
Beyond your own WeChat Official Account, articles mined in VoiceDrop can be shared to the VD Community — a public space of creators who also love turning speech into writing. Share a finished article, and others can read it, like it, toss coins, or even record a response article; you can read theirs and encourage them back.
The community is right at the top of the main screen: two tabs, "My Recordings" and "VD Community" — tap the latter to browse what everyone has shared. The list is sorted newest-first by share time (with some ranking based on your interests); pull down to refresh.
Every shared article gets a public link (like /voicedrop/<share-id>) you can forward to WeChat, X, and elsewhere — recipients read it right in the browser, no app install needed.
The community follows a set of Community Guidelines with zero tolerance for objectionable content: sexual content, violence, hate, harassment, illegal content, self-harm. You can report inappropriate content or block users you don't want to see at any time; reported content is taken down immediately and goes to human review.
How to use it
Share an article to the community
- Open a recording that has produced an article and go to its detail page (only recordings with a mined article can be shared).
- Tap the ⋯ menu in the top right and turn on the "Visible in VD Community" switch.
- If it's your first time posting, the "Community Guidelines" appear first — read them and tap "Agree & Publish".
- First-time sharing also requires Sign in with Apple (community posts need a real, accountable identity; the app prompts the login automatically).
- On success you'll see "Now visible in VD Community". If the article trips the sensitive-word filter, the share is rejected.
Browse and read
- Switch to the "VD Community" tab at the top of the main screen to see all shared cards; pull down to refresh.
- Tap any card to open the article. Shares with multiple articles show switchable title chips at the top.
Interacting on someone's article (detail page top bar and ⋯ menu)
- Tap ❤️ (heart) to like the article.
- Tap ⚡ (lightning circle) to "toss a coin" in encouragement.
- Tap ⋯ → "Write a response" and just record a voice memo; when you stop, the AI mines it into an article, published automatically as a "response" attached under the original.
- Tap ⋯ → "Share" to forward the article to WeChat, X, and more (a public link card is generated).
Withdraw a shared article (two ways, either works)
- Return to that recording's detail page, tap ⋯, and turn off "Visible in VD Community" — you'll see "Hidden from VD Community".
- Or find your own post in the "VD Community" list, choose "Remove from community", and confirm "Remove".
- Both only take it off the community — your original article is untouched, and you can share it again anytime.
Report / block
- On an article's detail page, tap ⋯ → "Report" → "Report & take down" — the article is immediately removed from the community pending review.
- Tap ⋯ → "Block this user" to stop seeing their community content; unblock anytime in "Settings → Blocked Users".
Details & rules
- What can be shared: only recordings that have produced an article. What goes public is the finished article and the scene photos in its body — the original audio recording is never made public.
- Community Guidelines (must agree once before your first post): you are responsible for what you publish and must have the right to publish it; sexual or explicit content, graphic violence, hate or discrimination, harassment or bullying, illegal content, and self-harm are strictly forbidden — zero tolerance.
- Posting requires Sign in with Apple: sharing, withdrawing, coin-tossing — all "write" actions require an accountable Apple identity. If you're anonymous, the app guides you through login and then retries automatically.
- How coin-tossing works: one toss gives the author 2 coins and you 0.5 coins, instantly converted to credits at the current coin rate. One toss per article per person; you can't tip your own article; tossing again shows "already tossed on this one".
- Daily tipping cap: when the day's credit pool is exhausted, you'll see "today's credit pool is used up — come back tomorrow".
- Responses: writing a response is just recording another memo that gets mined into an article, automatically linked to the original. The original shows "N responses" underneath, each openable.
- What reporting does: reported content is removed from the community immediately (you won't see it again either) and handled by human review within 24 hours. Repeat or serious offenders get removed.
- Blocking is local: blocking takes effect only on your own device (filtering by author nickname), the other person isn't notified, and you can undo it anytime in "Settings → Blocked Users".
- Contact & complaints: to reach us about content, email jianshuo@hotmail.com.
FAQ
Why does sharing to the community require Sign in with Apple?
If I withdraw from the community, is my original article deleted?
What is "coin-tossing"? What do coins do?
What do I do about inappropriate content?
Credits
Credits are VoiceDrop's billing currency — transcription, article mining, voice edits, and image generation all settle in credits.
Credits are VoiceDrop's unified billing unit. Everything AI-powered you do in the app — transcribing recordings, "mining" dictation into articles, editing articles by voice over and over, generating cartoon explainers — deducts from your credit balance by actual usage.
New users don't need to pay first. Signing up grants a one-time gift of 200 credits, valid for 365 days (one year) — plenty for everyday dictation and mining for a long while. You can keep earning credits through promotional gifts and community coin-tossing.
Your balance and every transaction are recorded on the Settings → Credits page, fully transparent: the big number on top is your remaining credits and roughly how many articles they can mine, with "Credit sources" and "History" sections below — the whole story at a glance. When your balance runs dry, credit-consuming actions are temporarily blocked with a prompt; everything already produced is unaffected.
How to use it
- Open Settings and tap "Credits".
- The dark card on top shows your remaining credits, with an "≈ N articles" estimate next to it (at roughly 9 credits per article), and a line below reading "total granted X · used Y".
- Below that, "Credit sources" groups everything you've received by source — like "sign-up gift", "promotional gift", "monthly allowance".
- Further down, "History" is the transaction ledger, newest first. Green "+" entries are income (sign-up gift, coins received); "−" entries are spending (mining, transcription, voice edits, image edits), each timestamped.
- In normal use, no action is needed — recordings upload, the system transcribes, mines, and bills automatically. You just talk.
Details & rules
- New-user gift: 200 credits on sign-up, valid for 365 days.
- Article mining: billed by the AI tokens actually consumed — roughly 9 credits per article (the app's "articles remaining" estimate uses 9 credits/article).
- Transcription: billed by recording length at ¥0.8/hour (about 18 credits/hour).
- Voice edits: billed by actual usage per edit; each article allows up to 100 edits, after which you'll see "this article has reached its edit limit (100)".
- Image edits / cartoon explainers: a flat 1.8 credits per image. When short, you'll see "not enough credits — one image costs 1.8 credits, please top up".
- Style distillation and Xiaohongshu copy: also deducted by actual AI usage.
- Recordings cap at 3 hours — longer ones aren't processed (marked too long).
- Community coin rewards: when you share an article to the community and someone tosses a coin, both sides receive credits (the author gets 2 coins per toss, the tosser 0.5, converted to credits at the current coin rate). These reward credits expire in 90 days. The daily pool is 2,000 credits, and a single coin converts to at most 200 credits. Repeated tosses from the same person to the same author taper off (2nd toss at 70%, 3rd and later at 50%) to prevent farming.
- Referral rewards: none at the moment. The ways to earn credits are the sign-up gift, promotional gifts, and community coins.
- Monthly subscription: the card shows ¥19.9/month, 200 credits per month, reset at month end, cancel anytime — but it's marked "coming soon" and not yet available; tapping it says it's still in development.
- When credits run out: transcription/mining of new recordings is paused and flagged (they can be processed once you have credits again), and voice edits show "not enough credits to keep editing". Finished articles and history are never lost.
- Pricing basis: internally, 23 credits = ¥1.
FAQ
How many credits do new users get, and how long do they last?
How much does mining an article or generating an image cost?
Besides topping up, how can I get credits for free?
What happens when I run out of credits? Do I lose my articles?
Developer API & Account Sign-in
VoiceDrop's account system (anonymous tokens / Sign in with Apple / 6+4 device pairing) and the open HTTP API for power users.
VoiceDrop's backend is a set of open APIs running on Cloudflare — your recordings, articles, photos, and writing style are all readable and writable with a single credential. If you just use the app normally to record and read articles, you can skip this whole section: login, tokens, and credits are all handled for you.
This section is for two kinds of people:
- People who want to sign in on a computer, or approve a computer login from their phone — understand how VoiceDrop accounts work (anonymous identity, Sign in with Apple, 6+4 phone device pairing).
- Power users who want to write their own scripts / clients against the API — VoiceDrop exposes articles, recordings, photos, mining, credits, community, and writing style all as HTTP endpoints; one token goes everywhere.
The data flow in one sentence: you upload an .m4a recording, the server mines it automatically (transcription first, then AI writing) into articles, and the client reads those articles back to share, publish to WeChat, or post to the community. Every step has a corresponding endpoint.
How to use it
1. How accounts work
VoiceDrop doesn't force registration. Its account system has three layers, all pointing to the same copy of your data:
- Anonymous identity (anon token): the first time you use the app, it locally generates a high-entropy token starting with
anon_— that's your identity. Your data lives in an isolated space that belongs only to you; nobody else can cross into it. You can copy this token from the app's Settings and use it to call the API elsewhere. - Sign in with Apple: signing in with Apple in the app simply binds your Apple ID to your current anonymous identity — your data doesn't move; it's still the same copy. Once bound, you can recover the account on a new device with the same Apple ID. Posting to the community requires Apple sign-in first.
- 6+4 phone device pairing: for signing in to your VoiceDrop account on a computer. Instead of hand-copying the long token, approve it from your phone (see below).
2. Approve a computer login from your phone (6+4 device pairing)
This is for when you want to log in on a computer (say, with a command-line tool) without hand-copying the token. The computer plays the "new device", the phone plays the "old device" — pair once, and the identity lands safely on the computer:
- Open the phone app's Settings → Account, which shows a 6-digit hexadecimal short ID.
- Start pairing on the computer and enter that 6-digit code.
- The phone immediately shows a 4-digit numeric code (with an "It's me / Not me" confirmation).
- Type the 4-digit code back on the computer — pairing complete, and the identity is now stored locally on the computer.
Prerequisites: the phone is online, the app is in the foreground, and it's signed in to the target account.
3. Calling the API yourself
- First get a token: copy the anon token from the app's Settings, run the 6+4 pairing above, or exchange an Apple sign-in for a session.
- Send
Authorization: Bearer <your token>on every request (the Files API also accepts a?token=query parameter). - Call the endpoints you need: list articles, read full text, upload recordings to trigger mining, check your credit balance, share, publish to WeChat, browse the community, and more.
- For the complete endpoint list, request examples, and response fields, see the official developer docs (linked under "Details" below).
Details & rules
Three backend services (Base URLs)
- Files API —
https://jianshuo.dev/files/api/: accounts, recording / article files, photos, sharing, WeChat, community, Apple sign-in. The vast majority of calls go here. - Agent Worker —
https://jianshuo.dev/agent/: triggering mining, live status push, voice editing, credit balance, 6+4 device pairing. - Reco Worker —
https://jianshuo.dev/reco/: community feed ranking and interaction reporting. Can be unplugged anytime; if it's down, the main flow is unaffected.
Three kinds of credentials
- anon token: a high-entropy string of 20+ characters starting with
anon_, generated once by the client and used long-term. Its scope (isolated data space) isusers/anon-<hash>/. - session JWT: obtained via
POST /files/api/auth/applewith an Apple identityToken, valid for 365 days. anon and session resolve to the same user and the same data. - 24-hour read-only temp token: issued by
GET /files/api/token/articles, withexpires_inof 86400 seconds. It can onlylist/download— for handing your article list read-only to external tools. The Agent and Reco services do not accept this read-only token.
Permission boundaries
- Your scope decides whose data you can see. All relative paths are automatically prefixed with your scope — there's no crossing out (
..and absolute paths are rejected outright). - Most endpoints accept any valid token; only community writes (share / withdraw) require an Apple-signed session, otherwise you get
403 needs_apple_signin. - The only public, token-free endpoints are: public photos
photo/*, the WeChat cover galleryasset/wechat-covers/*, the Apple sign-in exchangeauth/apple, and share-page HTML.
Endpoint groups (full list in the developer docs)
- Articles: list, read full text, write (versioned, with undo / redo), delete, sidecar flags (SRT / no-speech / out-of-credits), tags.
- Recordings: list, upload (upload auto-triggers mining), download.
- Photos: list, upload, download (private scoped, or public key).
- Mining:
POST /agent/mine/triggerprocesses pending recordings (the server also auto-sweeps every 6 hours). - Credits:
GET /agent/usage/balancefor the balance,GET /agent/usage/ledgerfor the transaction ledger (priced at 23 credits = ¥1). - Community: feed, read one, reply, share / withdraw (requires Apple sign-in), report.
- Writing style: read / write (versioned), style-learning corpus collection, server-side distillation, re-mining a single article with a chosen style.
Hard limits on 6+4 device pairing
- The pairing code is valid for 2 minutes (start over after a timeout).
- The 4-digit code allows at most 5 attempts, with a remaining-attempts hint on each miss; exhausting them voids the pairing.
- A 6-digit prefix matches at most 10 candidate accounts.
- Tapping "Not me" on the phone cancels the pairing immediately.
Security note: what 6+4 pairing hands over is the account's full identity key, not a separately revocable sub-token. Whoever holds the local credential file holds full control of the account — guard it carefully, and never commit or sync it anywhere others can read.
Developer documentation links
- Developer home:
/voicedrop/en/developer/ - HTTP API reference (complete endpoints + request examples):
/voicedrop/en/developer/api.html - Claude Code command-line tool guide:
/voicedrop/en/developer/wjs-voicedrop.html