AI proof of concept for companies in Germany
Before your company sets a budget for an AI idea, a proof of concept shows whether the idea works. It tests one task on your own data against a pass mark we agree on before I build anything, and it ends with a measured result. I take on freelance and interim engagements, remote from Frankfurt, with on-site days when the work needs them. Let's talk about your needs in a free 30-minute video call.
What I test for your company
-
Experience with systems and their users
For more than twenty years I have designed systems for the people who use them and challenged concepts that only worked on paper. An AI test faces the same question: does it help your staff?
-
A pass mark set in advance
Before I build, we write down what the AI must reach, for example 95 of 100 delivery notes read correctly. The test ends with a number against that mark.
-
Results measured by a script
In a client's project in June 2026 an AI agent did its job correctly and reported 10,481 records instead of 1,832. Since then a separate script measures each result.
-
Coding agents under written rules
I build with Claude and Codex under written rules. For a client this work has produced more than 300 API endpoints and more than 8,000 commits.
-
Knowledge for your team
I teach practical work with Claude at Claude Hacker House. During the test your staff learn how the AI step works and where it fails.
-
A decision you can take
You receive the measured result and the list of cases the AI got wrong. Whether the idea goes further stays your decision.
What an AI proof of concept answers
An AI proof of concept answers one question: does an AI do a given task on your data well enough to be worth a budget? It is a small, working test with a number at the end.
Management has seen AI demos, and staff use ChatGPT on their own. Nobody in the company knows yet whether the AI reads your delivery notes or your support tickets as reliably as the text in the demo. A proof of concept finds out on real cases from your company, before anyone plans a project around the idea.
How I build and measure the test
The test starts with the answers. Your department picks the cases, and a person who knows the work writes the right answer for each one before the AI sees them. Then I build the AI step on a development system, with Claude or Codex as coding agents, and run it on the same cases. A script compares each answer of the AI with the answer your department wrote and counts the hits.
I measure with a script because the report of an AI can be wrong while its work is right. In June 2026 an agent in a client’s project assigned 1,831 of 1,832 records correctly and then reported 10,481 records, because it had counted a wider area than the task named. Details: Running a company of AI agents with Paperclip.
The AI in the test works under the rules I use for coding agents. It gets the rights of a new colleague on the first day: it reads copies of data and cannot change anything in your live systems. Details: AI agents under the same rules as humans.
The test sheet: seven decisions before the build
Fill in the sheet with your team before our first call, or we fill it in together in the assessment. The examples show one possible task.
| Decision | Example | Who decides |
|---|---|---|
| The task | Read the order number and the delivery date from supplier delivery notes | The department that does the work today |
| The cases | 100 delivery notes from the last quarter, bad scans included | The department |
| The right answers | Written by a clerk for each case before the AI runs | The department |
| The pass mark | 95 of 100 correct, and no wrong delivery date passed on without a warning | Management |
| The data rules | Copies without personal data, on a development system | Your data protection officer |
| The model | A cloud model through its API, or an open-weight model on your own server | IT |
| The next step after a pass | A production version with user rights and logging, or a second task | Management |
A test without the third row measures nothing: without the right answers written down in advance, the result is an impression.
Who I am
I am Michael Wutzke from Frankfurt. I have worked in IT and media for more than twenty years, among other roles as CIO at the Frankfurt-based company Blocksize Capital. Today I work as an AI engineer on a client’s platform with Claude and Codex, and at Claude Hacker House I teach practical work with Claude. I am also interested in open-source AI models that a company runs on its own servers. Details: Career stages.
How an engagement runs
-
Free video call
In 30 minutes we talk about the task the AI should take over and the data it would need.
-
Assessment together
We choose one task, look at sample documents or records and agree on the pass mark. Your department names the person who writes the right answers.
-
Quote and order
The assessment ends with a quote. The build starts when you accept it.
-
Build and measure
I build the test on a development system and run it on the agreed cases. A script compares each answer of the AI with the answer your department wrote.
-
Result and decision
We go through the result and the failed cases together. You decide whether the idea moves on to a production version.
Questions companies ask
What is an AI proof of concept?
A working test of one AI task on a company's own data. It answers one question: does the AI do this task well enough to justify a budget? A clickable prototype shows screens for a decision, and an MVP puts a first version in front of real users.
Which AI model does the test use?
The one your data rules allow: Claude or ChatGPT through a business plan or an API, or an open-weight model on a server of your own. The page Self-hosted LLMs such as DeepSeek covers the last case.
May our data go into the test?
Your contracts and your data protection officer decide that. We agree in the assessment which data the test uses. Where names play no part in the task, the test runs on copies without them.
What happens after a successful test?
The code of the test belongs to your company. A production version needs more: user rights, logging, error messages that reach a person, and someone who confirms what the AI changes. I can build it, or your team takes it over.
What if the AI misses the pass mark?
Then your company has spent the budget of a test. The failed cases show where the AI fell short, and whether a narrower task could pass.
Which engagements do you take on?
Freelance and interim engagements, part time or full time. I work remote from Frankfurt and come on site in Frankfurt Rhine-Main or elsewhere in Germany when the work needs it.
Details on the work behind this page
-
AI agents under the same rules as humans
Limits at the resource, a blocking check before each tool call, and approvals that only a person gives.
-
Running a company of AI agents with Paperclip
An org chart of AI agents on the open-source control plane Paperclip, June 2026: the delegation rule, the human gate and four lessons from the first jobs.
-
Product conception
From the first concept to the roadmap: advice on digital products, based on more than twenty years of product and IT projects.
-
Teaching and certifications
Where I teach, my certificates with the issuer of each, and my education.
Your AI proof of concept developer in Germany
I am Michael Wutzke, a freelance AI developer in Germany, based in Frankfurt. In a free 30-minute video call we talk about your AI idea and your needs, and you learn which task a proof of concept should test first and how we would measure it.
Searches this page answers
- AI proof of concept
- AI proof of concept (POC)
- AI proof of concept development services
- generative AI proof of concept
- agentic AI proof of concept
- AI proof of concept failure
- AI POC
- AI POC development
- AI POC development services
- AI POC and MVP development services
- AI POC examples
- agentic AI POC examples
- AI POC to production
- agentic AI POC to production
- AI prototype
- AI prototyping
- AI prototype development
- AI prototyping services
- AI prototyping engineer
- rapid AI prototyping
- AI agent prototype
- AI agent prototype to production
- AI feasibility study
- AI feasibility assessment
- AI feasibility analysis
- AI pilot project
- AI use case validation
- agentic AI use case validation
- AI use cases for small businesses
- best AI use cases for small business
- freelance AI developer
- hire freelance AI developer
- freelance AI app developer
- prototype with Claude
- Claude AI prototype
- POC vs MVP
- proof of concept vs MVP
- difference between proof of concept and MVP