Computer Vision · Physical AI · Robotics

Vision systems that do not only see - they understand.

What we deliver.

Our engineers bring the latest generation of vision models into medical devices and robotic systems. Instead of a single object detection you get a layer that understands the situation and can propose a safe course of action.

A generational shift

From recognising objects to understanding situations

A classic vision system answers one question: what is in the frame. That is enough on a production line, where everything is predictable. It is not enough in an operating room, or anywhere the situation changes from second to second.

Classic computer vision
  1. camera
  2. detection
  3. classification
  4. control algorithm
“The cup is 20 cm in front of the robot.”

Precise, fast and completely blind to context. Every exception has to be anticipated and programmed in advance.

Vision-Language-Action
  1. image + language
  2. scene understanding
  3. reasoning
  4. action plan
“This is a cup. It stands on the edge of the table. If I grip it from this side I may knock it over — I will approach from the other side.”

The model combines the image with knowledge of the world and of physics. It does not need every exception described — it can assess a situation nobody programmed in advance.

The foundation

What we build on: foundation models for Physical AI

NVIDIA calls this area Physical AI, and it does not ship a single model but a complete toolchain: from world understanding, through simulation and training, to on-device deployment. A company building a robot therefore no longer has to create the whole “brain” from scratch — it can take an existing model and tune it to its own machine.

  • Scene understanding

    NVIDIA Cosmos / Cosmos-Reason

    A model that analyses images and video in terms of 3D space, relations between objects, motion over time and physics. It answers not only “what do I see?”, but also “what is happening here and what happens next?”.

  • Vision + language → action

    NVIDIA Isaac GR00T (VLA)

    A Vision-Language-Action model: from an image and an instruction given in plain language it generates a sequence of robot movements. Released with weights and code under Apache 2.0, so it can be fine-tuned to a specific machine and a specific task.

  • Simulation and data

    Isaac Sim / Isaac Lab

    An environment where a robot learns in a virtual world before it touches anything real. It generates thousands of realistic training scenes where real data is scarce — and in medicine it always is.

  • On-device deployment

    TensorRT / ONNX / Jetson

    The path from a research model to a palm-sized module working next to the operating table. Model compression and optimisation so that a decision is made in milliseconds and locally — without sending imagery to the cloud.

We are not tied to a single vendor. The NVIDIA family is currently the most complete one for robotics, but the choice of model always follows the task: sometimes the right answer is a large vision-language model, and sometimes a small specialised network running in a few milliseconds on modest hardware. We start from the requirements, not from a model name.

Architecture

What such a layer looks like in practice

Artificial intelligence never drives the motors directly. The model proposes an action, and a separate, deterministic safety layer decides whether it may be executed. That separation is essential wherever a human being stands on the other side of the machine.

AI MODELSDETERMINISTIC LAYERCameraimage / videoPerceptionwhat is in frameUnderstandingspace and relationsPlanproposed movementSafetylimits and checksControllertrajectoryRobotmotionOperator: oversight and stop at any momentfeedback — a new view of the scene
Team capabilities

What our engineers can actually deliver

We work at the intersection of medical imaging, artificial intelligence and hardware. The same team that built three-dimensional visualisation of patient data for the operating room now designs perception layers for robotic systems.

  • Image and video perception

    Real-time detection, segmentation and tracking of objects — instruments, anatomical structures, people and obstacles in the robot's working field.

  • 3D spatial understanding

    Reconstructing scene geometry, estimating object position and orientation, and aligning the camera view with patient imaging data (CT, MR) in a single coordinate system.

  • Vision-language and VLA models

    Selecting, fine-tuning and evaluating models such as Cosmos-Reason or GR00T on the client's own data — so that the model understands the specific instruments, procedures and vocabulary of the team.

  • Synthetic data and sim-to-real

    Building simulation scenes, generating training data and transferring learned behaviour from simulation to the physical machine, with an honest validation of the gap.

  • Optimisation and edge deployment

    Quantisation, export to ONNX/TensorRT and deployment on industrial computers or Jetson modules, with latency and power measured on the target hardware.

  • Safety layer and validation

    Deterministic control over what the model is allowed to execute, decision logging, reproducible testing and documentation prepared with medical device requirements in mind.

The governing rule

AI proposes. The layer that can be verified decides

The model may say: “move the instrument 4 mm to the left”. Before anything moves, that proposal passes through a set of simple, predictable rules that can be tested, documented and presented to a certification body.

As a result, an advanced model never has to be treated as a black box holding a scalpel. It is a broad-minded advisor whose conclusions are checked by code written to medical device rigour.

Questions the safety layer answers
  • Is the movement inside the permitted workspace?
  • Does it avoid collision with the patient, the team or another instrument?
  • Do speed and force stay below the agreed limits?
  • Can the operator take over or stop the machine at any moment?
  • Was the decision logged in a way that allows it to be reconstructed later?

Every answer is deterministic: with the same inputs, the outcome is always the same.

Applications

Where such systems make sense today

  • 01

    Surgical robotics

    A perception layer that sees the instrument, the tissue and their relative position, and that can warn before a movement becomes risky.

  • 02

    Intra-procedural navigation

    Combining the camera view with a three-dimensional reconstruction of patient data — the direction we develop in CarnaLife Holo and MedNav.

  • 03

    Laboratory automation

    Robots handling samples and equipment in a changing environment, where not every possible situation can be programmed in advance.

  • 04

    Assistive robots and rehabilitation

    Systems that observe the patient's movement and adapt their own behaviour to what is actually happening at a given moment.

  • 05

    In-hospital logistics

    Vehicles and manipulators moving among people, for which understanding the intent of their surroundings is a safety requirement.

  • 06

    Quality control and documentation

    Automatic analysis of procedure or production-line recordings: what happened, in what order, and whether it followed the procedure.

How we work

From an idea to a working device in four steps

  1. 01.

    Feasibility study

    We start by asking whether the problem suits a vision model at all. We review the available data, lighting conditions, latency requirements and hardware constraints, and then say plainly what is realistic and what is not.

  2. 02.

    Prototype on your data

    We build a working proof of concept on real recordings and judge it by numbers rather than impressions: detection performance, spatial accuracy and reaction time.

  3. 03.

    Integration and simulation testing

    We connect the model to the robot control through the safety layer and test it in simulation — including scenarios that must never be triggered deliberately in the real world.

  4. 04.

    Deployment and maintenance

    We optimise the model for the target hardware, hand over documentation and tests, and then keep improving the solution as new operational data arrives.

You have a machine that should see more?

Tell us what your robot or device should understand. We will start with a short conversation and an honest feasibility assessment — including when the answer is “not yet”.

Contact

Address

Aleja Juliusza Słowackiego 6/10

30-037 Kraków

Poland

Follow us

NIP: 6631868308

REGON: 260637552

KRS: 0000443282

MedApp Resources

Login
2026 © Medapp S.A.Terms & ConditionsPrivacy Policy
EU banner