What is UD's AI vs Human challenge, and what does it cost?
UD's AI vs Human challenge is a free browser test with 20 industry arenas where you answer timed workplace judgement questions against an AI opponent's answer key. It costs HK$0, requires no registration, no email and no card, and one arena takes about five minutes.
I opened it expecting a marketing quiz. What loaded was a 30-second countdown and a staffing problem: 150 CVs, three marketing supervisor roles, a HR team that needs three weeks for first-pass screening, and interviews starting in two weeks. Four plausible answers. Clock running.
That is a better test than it sounds, and also a narrower one than the framing suggests. This review covers what the test actually measures, what it does not, and whether it is worth the five minutes for someone who already uses AI daily.
You can open it directly at the AI vs Human arena.
What actually happens when you open an arena?
Each arena runs 10 multiple-choice questions with a 30-second timer on each and four options per question. A running score sits at the top of the screen next to the question counter. There is no landing form, no account creation and no payment step between the arena link and question one.
The questions are scenario-based operational decisions, not knowledge recall. The HR recruitment arena's first question is a screening bottleneck. The options range from mandatory overtime, to an AI screening pass with human final confirmation, to reading only the first 50 CVs, to outsourcing the first pass.
The verified format, as of 5 August 2026
--- Cost: HK$0. No free tier limits, because there is no paid tier of the test.
--- Registration: none. No card, no email, no account.
--- Length: 10 questions per arena, 30 seconds each, roughly five minutes including reading time.
--- Arenas: 20, each with a named AI counterpart role.
--- Output: an on-screen score out of 10. No PDF, no certificate, no email report.
--- Interface language: Traditional Chinese only, written in Hong Kong colloquial style. No English or Simplified Chinese version.
Which of the 20 arenas should you pick?
Pick the arena closest to a decision you actually make weekly, not the one closest to your job title. The 20 arenas cover catering food safety, retail inventory, beauty skin analysis, property matching, education diagnosis, medical triage, logistics dispatch, accounting tax, fitness coaching, hotel concierge, legal contract, insurance underwriting, renovation quotation, auto repair, pet care, wedding planning, HR recruitment, food delivery, construction safety and florist consulting.
Notice what is missing. There is no arena for content marketing, SEO, design, social media or software work.
If you are a marketer, a content creator or a freelancer, no arena matches your title. That is not a dealbreaker, but it changes how you should use the test.
Practical mapping for practitioners
--- If you triage inbound requests: run HR recruitment or medical triage. Both are prioritisation-under-time-pressure problems.
--- If you scope and quote work: run renovation quotation. It is the closest thing to a freelance pricing decision on the list.
--- If you manage stock, budgets or campaign spend: run retail inventory. The structure is forecast against constraint.
--- If you write contracts or client agreements: run legal contract.
What does the test actually measure, and what does it not?
It measures whether you reach a defensible operational decision in 30 seconds when four options all look reasonable. It does not measure your prompting skill, your AI output quality, or whether an AI could do your job. Those are three different things and the framing invites you to confuse them.
Be clear about the mechanics before you read anything into your score. The arena presents a fixed question set with a scored answer key, so treat the AI side as a designed benchmark rather than a live model response. That makes it a decision-quality quiz, not a model evaluation.
Honest limitations, named
--- It is not a model benchmark. If you want to know whether GPT-5.6, Claude or Gemini writes a better first draft, this test cannot tell you. For that, run the same brief through two tiers and compare, as described in this guide to routing work across model tiers.
--- The 30-second timer rewards pattern recognition. A low score often means you were reading carefully, which is the correct behaviour in real work and the wrong behaviour here.
--- The answers have a point of view. Options that pair AI with human final confirmation tend to be the defensible ones. That is a reasonable position and it is also UD's commercial position, so read the test as a framing exercise rather than a neutral assessment.
--- Nothing is portable. No certificate, no PDF, no shareable result. You cannot put this on a CV or in a performance review.
--- Chinese only. If you work in English or need to share it with a non-Cantonese colleague, the language is a hard blocker.
--- No published price for the actual product. The test is free, but UD does not publish a fixed price for AI Staff deployment. You reach that number through a consultation, not a pricing page.
What happens after you finish, and is a card ever required?
You see your score out of 10 and nothing else happens. No card is requested at any point, no account is created, and no email is captured, so there is no drip sequence to unsubscribe from. Follow-up only happens if you initiate it.
The contact routes are published on the same pages: WhatsApp on (852) 9696 7545, phone on (852) 2554 7545, and email to sales@ud.hk.
That is unusually low friction for a lead-generation tool, and it is the main reason the five minutes is defensible. There is no cost, no data trade and no obligation.
The trade-off is on the other side. Because nothing is captured, nothing is saved. Run the arena twice and the first score is gone.
How do you get more out of it than a score?
Screenshot each question before the timer runs out, then rebuild the same decisions as prompts afterwards. The value is not in beating the answer key. It is in discovering which of your recurring decisions have a defensible answer an AI can already reach, and which genuinely need your context.
Here is the exercise that turns a five-minute quiz into something reusable.
Copy this prompt after you finish an arena:
I just answered a timed multiple-choice question about an operational decision in my field. Here is the scenario and the four options.
[paste the scenario and all four options]
Do four things, in order.
1. Tell me which option you would choose and why, in under 80 words. Then tell me which option is the most commonly recommended answer in professional practice, and say clearly if that differs from your own choice.
2. Name the piece of information missing from the scenario that would most change your answer. Just one, the highest-impact one.
3. Rewrite the scenario as a decision I could delegate to an AI reliably every week. State exactly what inputs the AI would need, what output format it should return, and what condition should send the decision back to me instead.
4. Tell me honestly whether this decision should be delegated at all, and name the specific risk if it is delegated badly.
Do not flatter my original answer. If my reasoning has a gap, say so directly.
Run that on two or three questions and you end up with something more useful than a score: a short list of decisions in your own week that are already automatable, and a shorter list that is not.
Who should skip this test?
Skip it if you need a rigorous, citable assessment, if you need an English interface, or if your question is about model output quality rather than decision quality. Three specific profiles should spend the five minutes elsewhere.
Skip if you need a credential. There is no certificate and no reimbursement pathway. If your goal is something you can show an employer, a paid course is the honest comparison, and we worked through the real Hong Kong costs in Is a Paid AI Course Worth It in 2026?
Skip if you are benchmarking models. Use a side-by-side test on your own real brief instead. That takes 20 minutes and tells you something this test structurally cannot.
Skip if the language is a barrier. The interface is Hong Kong Traditional Chinese only, and a timed test in a second language measures your reading speed, not your judgement.
Everyone else: it costs HK$0, takes five minutes, asks for nothing, and reliably surfaces at least one routine decision you had never thought of delegating.
The verdict
As a benchmark, this is a light tool. As a prompt for a specific and uncomfortable question, it works well: which of the decisions you make on autopilot already has a defensible answer that does not need you?
The honest recommendation is to treat the score as noise and the questions as raw material. Five minutes in, one prompt afterwards, and you have a shortlist of work worth automating.
That is the part no tool can do for you, and it is the part worth having help with.
We understand AI. We understand you better. With UD by your side, AI doesn't feel cold. Twenty-eight years in Hong Kong technology has taught us that the hard part was never the tool.
Reviewed by the UD AI team. Product facts verified against the live pages on 5 August 2026.
Take the Challenge, Then Build the Workflow
The test is free and takes five minutes. Turning what it reveals into a workflow that runs every day is the harder part, and we'll walk you through every step, from choosing the right task to configuring it and putting it into production.