AI-era talent
evaluated in a new way
Don't measure tool fluency — measure the thinking: how someone explains a problem to AI, validates the output, and chooses the right tool. Beyond the deliverable, “how they solved it” is captured automatically, so reviewers can see the whole process at a glance.

Talent in the AI era
delivers different results under the same conditions
Handing everyone the same AI doesn't make the output the same. What sets them apart isn't the tool — it's the competency to operate it.
So, what are you looking at
to hire AI talent today?
What makes a good AI hire, what 'using AI well' actually means, how much AI literacy the people evaluating themselves should have — none of it is clear.
“No AI hiring profile”
Because each company and role needs different AI abilities, there's no agreed-upon, universal definition of an AI hire.
“No yardstick for 'good with AI'”
Shipping code fast isn't the same as using AI well. You need a measure of what to look for, and how far it should go.
“I'm a reviewer, but I barely know AI”
If the people evaluating aren't comfortable with AI tools and environments, they can't read the real gap between candidates.
Defining AI talent, setting the bar, automating assessment —
all in one module
Probe gives recruiters an AI hiring standard, co-authors the problems with them, and runs an automated evaluation.
Usecase


The future of hiring evaluation,
built with an AI-first company
Probe is in production as Krafton's standard hiring assessment. From researchers and engineers to non-developer and back-office roles, Probe delivers trusted tests and assessment reports aligned to Krafton's Vision & Value.
AI-native hiring know-how
Krafton's criteria and methods for hiring at the AI frontier — packaged in.
In production, expanding
Krafton itself is rolling AI capability assessment across more and more roles.
Our Standard
We define AI-native competency,
and build a new standard of evaluation
The AI-native competency companies want is the ability to produce great work by operating AI's strengths and limits within constrained resources. We evaluate it along three axes: commanding the artifact (technique), steering the process (intent), and understanding the result (cognition).
What is AI-native competency?
- 01
Technical control (Technique)
Is what the AI produced well put together under your own command — running cleanly, maintainably structured, and free of leaked secrets?
- 02
Intent management
Do you steer the AI — defining goals and constraints, decomposing and delegating work, and controlling results through verification loops?
- 03
Cognition management
Can you explain the AI's work grounded in your own artifact — why this structure, how it runs, and where it breaks?
Evaluation areas
We evaluate on technique, intent, and cognition
Each competency is scored against a detailed rubric. Every criterion gets a score, with met and unmet examples presented alongside the prompts and source code behind them — giving evaluators a signal they can trust.
Technical control — scored on the artifact
Applied to every answer: no contradictions with the actual submission · grounded in their own artifact, not generalities · nothing fabricated · reasoning from evidence, not authority.
How it works
Candidates reason through the problem,
and companies evaluate that reasoning
Candidates solve the task in a real environment, then answer questions about what they built. Both behavior (intent) and answers (cognition) are evaluated from real records.
- 1
Start in a real environment
Launch instantly in a browser IDE. Real APIs and real compute — not mocks — within a time and token budget.
- 2
Artifact & process, auto-collected
Source code, along with AI collaboration records like the transcript, prompts, tool configuration, and verification runs, is collected automatically. Candidates never write a separate report.
- 3
Post-task Q&A
Post-task questions span the artifact, intent, and cognition, and candidates must be able to explain their own work.
- 4
Automated scoring · interview report
Auto-scored against per-competency rubrics, producing a discriminating per-candidate report for the interview panel. Every criterion defines its grey area and met/unmet examples, so scoring stays reproducible.
Signals that count as met
- Prompts with explicit done conditions & acceptance criteria
- Fail → fix → re-pass through the same verification
- Recorded rationale and rejected alternatives
- Intent docs a new session can inherit (CLAUDE.md)
- Answers grounded in their own artifact — matching what was submitted
Signals that count as unmet
- Wholesale delegation like "just make it work"
- Accepting AI's 'done' report without verification
- Answers filled with textbook generalities
- Explanations citing files or APIs that don't exist
- Signal drowned by tool & MCP sprawl
Same starting line for everyone
and we focus on the problem-solving process
We place no caps on model power or scope — candidates get the same freedom they'd have in real work.

