인공지능의 다음 경쟁은 더 자연스러운 문장을 만드는 능력이 아니라, 복잡한 현실을 이해하고 안전하게 움직이는 능력에서 벌어지고 있다. 보고, 듣고, 만지고, 판단한 뒤 실제 행동으로 이어지는 ‘물리적 AI’가 로봇·제조·물류·의료·돌봄을 하나의 기술 흐름으로 연결하기 시작했다. AI가 물리적 세계에 진입하는 순간, 성능의 기준도 정확한 답변에서 적절한 행동과 안전한 책임으로 바뀐다.
[Key Message]
* AI의 무대가 디지털 공간에서 물리적 세계로 확장되고 있다. 물리적 AI는 정보를 생성하는 데 그치지 않고 현실을 지각하고 판단하며 직접 행동한다.
* 지능은 알고리즘뿐 아니라 몸과 환경의 상호작용에서 형성된다. 로봇의 구조와 감각, 움직임은 지능을 구현하는 핵심 요소이며 몸 자체가 계산의 일부가 된다.
* 물리적 AI의 경쟁력은 시각·언어·행동을 연결하는 능력에 달려 있다. 세계 모델과 시뮬레이션, 촉각·힘·관절 데이터를 결합해야 낯선 환경에서도 적절한 행동을 선택할 수 있다.
* 공장·물류·의료·돌봄이 물리적 AI의 주요 적용 분야로 떠오르고 있다. 단순 반복 작업의 자동화를 넘어 환경 변화에 적응하고 인간의 판단과 행동을 보조하는 방향으로 발전한다.
* 행동하는 AI의 핵심 기준은 지능의 크기가 아니라 안전성과 책임성이다. 불확실할 때 멈추고 인간에게 통제권을 돌려주는 능력이 현실에서 신뢰받는 AI를 결정한다.
***
화면 속 지능에서 행동하는 지능으로
지난 수년간 인공지능의 발전은 대부분 디지털 공간에서 이루어졌다. 대규모 언어모델은 방대한 문서를 학습해 질문에 답하고, 생성형 AI는 문자·이미지·음악·영상·프로그램 코드를 만들어냈다. 예측 모델은 날씨와 금융시장, 질병 발생 가능성을 계산했고, 추천 알고리즘은 소비자의 다음 선택을 추정했다. 이들 시스템은 놀라운 능력을 보여주었지만, 활동 무대는 주로 화면과 서버 안에 머물렀다.
물리적 세계는 디지털 공간과 전혀 다른 조건을 지닌다. 문서에서 단어 하나를 잘못 선택하면 수정할 수 있지만, 로봇이 물건을 잘못 집으면 제품이 파손될 수 있다. 생성된 이미지의 손가락 모양이 어색한 것은 품질 문제로 끝날 수 있지만, 수술 로봇이나 돌봄 로봇의 손이 잘못 움직이면 사람의 안전이 위협받는다. 챗봇은 답변을 다시 생성할 수 있지만, 자율주행차는 갑자기 나타난 보행자 앞에서 수십 밀리초 안에 단 한 번의 행동을 선택해야 한다.
2026년 4월 『Nature Machine Intelligence』에 게재된 사설 「From Embodied Intelligence to Physical AI」는 이러한 전환에 주목했다. 사설이 던진 중심 질문은 분명했다. 인공지능이 세계를 예측하거나 시뮬레이션하고 추론하는 수준을 넘어, 실제 환경에서 지능적으로 행동하려면 무엇이 필요한가. 이는 단순히 기존 AI를 로봇에 탑재하는 문제가 아니었다. 지능이 어떻게 몸과 결합하고, 몸이 환경과 상호작용하며, 그 경험이 다시 학습과 판단을 변화시키는지를 묻는 과학적 과제였다.
‘물리적 AI’는 특정한 하나의 알고리즘이나 제품을 뜻하지 않는다. 현실을 감지하는 센서, 상황을 해석하는 인공지능 모델, 움직임을 설계하는 제어 시스템, 힘을 전달하는 로봇의 몸체가 하나의 순환 구조로 작동하는 기술 체계를 가리킨다. 카메라로 물체를 발견하는 것만으로는 부족하다. 그 물체가 무엇인지, 어디에 있는지, 어느 정도의 힘으로 잡아야 하는지, 움직였을 때 주변 환경이 어떻게 변할지를 판단해야 한다. 행동의 결과를 다시 감지하고, 예상과 다르면 즉시 계획을 수정해야 한다.
그동안 인공지능은 주로 세계에 관한 정보를 처리했다. 이제는 세계 안에서 행동하며 그 결과를 감당해야 한다. 물리적 AI의 부상은 로봇 기술의 진전만을 의미하지 않는다. 인공지능이 ‘말하는 도구’에서 ‘행동하는 주체’로 이동하고 있음을 보여주는 변화다.
몸이 지능의 일부가 되는 체화 지능
물리적 AI를 이해하려면 먼저 ‘체화 지능’의 의미를 살펴봐야 한다. 전통적인 인공지능 연구에서는 지능을 뇌나 컴퓨터 내부에서 이루어지는 정보처리 과정으로 보는 경향이 강했다. 외부 환경을 입력하면 내부 모델이 이를 계산하고, 정해진 결과를 출력한다는 구조였다. 이 관점에서 몸은 지능이 내린 명령을 수행하는 수동적인 장치에 가까웠다.
체화 지능은 이러한 구분에 의문을 제기한다. 지능은 뇌에만 들어 있는 것이 아니라 몸의 형태와 감각, 움직임, 환경과의 관계 속에서 형성된다는 관점이다. 인간이 컵을 집는 행동을 생각해보면 쉽게 이해할 수 있다. 사람은 컵의 위치만 계산하지 않는다. 손가락의 관절 구조, 피부에서 느껴지는 마찰과 압력, 컵의 무게, 물이 흔들리는 감각을 동시에 활용한다. 컵이 미끄러지면 시각적 분석이 끝나기를 기다리지 않고 손가락의 힘을 즉시 조절한다.
이처럼 몸은 단순한 명령 수행기가 아니라 계산의 일부다. 새의 날개와 물고기의 지느러미는 환경과 상호작용하며 복잡한 움직임을 자연스럽게 만들어낸다. 인간의 근육과 힘줄도 충격을 흡수하고 균형을 유지하는 데 필요한 계산 부담을 줄여준다. 로봇의 몸을 어떻게 설계하느냐에 따라 AI가 학습해야 할 문제의 난도도 달라진다. 유연한 관절이나 탄성 소재를 사용하면 충돌을 흡수하고 불규칙한 표면에 적응하기 쉬워진다. 몸의 물성이 제어 알고리즘의 일부 역할을 맡는 셈이다.
사설은 물리적 세계에서 행동하는 지능을 설명하기 위해 여러 학문적 흐름이 한 지점으로 모이고 있다고 보았다. 세계의 변화를 내부적으로 예측하는 모델, 환경과의 직접적인 상호작용을 강조한 행동 기반 로봇공학, 지각을 행동 가능성과 연결한 생태심리학, 몸의 구조와 물질 자체를 활용하는 소프트 로봇공학, 생물의 신경계와 진화를 모방하려는 접근이 서로 만나고 있었다.
다만 체화 지능과 물리적 AI는 완전히 같은 표현은 아니다. 체화 지능은 몸과 환경의 상호작용이 지능 형성에 어떤 역할을 하는지를 탐구하는 과학적·철학적 개념에 가깝다. 물리적 AI는 이러한 통찰을 활용해 현실에서 작동하는 시스템을 만들려는 공학적 흐름을 넓게 포괄한다. 가상공간의 아바타도 체화된 에이전트로 연구될 수 있지만, 물리적 AI는 실제 중력과 마찰, 충돌, 재료의 변형, 센서의 오차를 견뎌야 한다.
물리적 AI의 진전은 더 큰 언어모델 하나만으로 이루어지지 않는다. 지능의 일부가 몸에 있고, 또 다른 일부가 환경과의 관계 속에 존재한다는 인식이 로봇 설계와 AI 학습을 동시에 바꾸고 있다.
지각과 언어와 행동을 잇는 새로운 학습
기존 산업용 로봇은 정해진 환경에서 뛰어난 능력을 발휘했다. 자동차 공장의 로봇 팔은 동일한 위치로 들어오는 부품을 빠르고 정밀하게 용접했다. 그러나 부품의 위치가 조금만 달라지거나 예상하지 못한 장애물이 나타나면 작업을 중단해야 했다. 높은 정확성을 확보하는 대신 환경의 변화 가능성을 최대한 제거하는 방식이었다.
물리적 AI가 지향하는 로봇은 다르다. 사전에 모든 상황을 프로그램하지 않아도 사람의 지시를 이해하고, 처음 보는 물체와 공간에 적응하며, 행동 도중 발생한 변화에 대응해야 한다. “식탁을 정리해”라는 명령을 받았을 때 컵·접시·음식물·깨지기 쉬운 물건을 구분하고, 무엇을 먼저 옮길지 판단하며, 사람이 다가오면 동작을 늦추거나 멈춰야 한다.
이러한 능력의 중심에는 시각·언어·행동을 연결하는 모델이 있다. 시각언어행동 모델은 이미지와 문장을 이해하는 데 그치지 않고, 로봇이 수행할 동작을 출력한다. 카메라 영상으로 주변을 파악하고 자연어 지시에서 목표를 추출한 뒤, 팔과 손가락이 이동해야 할 방향이나 관절의 움직임을 결정한다. 기존 생성형 AI가 다음 단어를 예측했다면, 행동 모델은 다음 움직임을 예측한다.
그러나 현실의 행동은 토큰 생성보다 훨씬 까다롭다. 같은 컵이라도 유리·종이·금속에 따라 잡는 힘이 달라야 하고, 컵 안에 액체가 들어 있다면 기울기와 속도도 조절해야 한다. 책상 위에서 성공한 동작이 움직이는 차량이나 흔들리는 선박에서는 실패할 수 있다. 카메라에 잘 보이는 물체도 그림자나 반사광, 먼지, 가림 현상 때문에 다르게 인식될 수 있다.
이때 중요한 역할을 하는 것이 ‘세계 모델’이다. 세계 모델은 행동하기 전에 그 결과를 내부적으로 예상한다. 로봇이 상자를 밀면 어느 방향으로 움직일지, 문손잡이를 돌리면 문이 어떻게 열릴지, 빠르게 이동하면 물체가 넘어질 가능성이 있는지를 추정한다. 여러 행동의 결과를 미리 비교할 수 있다면 로봇은 시행착오를 줄이고 더 안전한 계획을 선택할 수 있다.
대규모 시뮬레이션과 합성 데이터도 물리적 AI의 핵심 기반이 됐다. 현실에서 로봇 한 대를 수천 번 넘어뜨리며 보행을 가르치는 방식은 시간과 비용이 많이 들고 장비도 손상시킨다. 반면 가상환경에서는 중력과 마찰, 조명, 지형, 물체의 배치를 바꾸면서 수많은 경험을 빠르게 생성할 수 있다. 디지털 트윈으로 공장이나 창고를 재현하면 실제 설비를 멈추지 않고도 다양한 작업과 위험 상황을 시험할 수 있다.
하지만 시뮬레이션은 현실의 근사치일 뿐이다. 가상공간에서 완벽하게 걷던 로봇도 실제 바닥의 미세한 미끄러움이나 관절의 마모, 센서 지연 때문에 넘어질 수 있다. 이른바 ‘시뮬레이션과 현실의 간극’을 줄이려면 가상 학습과 실제 경험을 반복적으로 연결해야 한다. 현실에서 수집한 데이터를 시뮬레이션에 반영하고, 시뮬레이션에서 배운 정책을 실제 환경에서 검증한 뒤 다시 수정하는 순환 구조가 필요하다.
물리적 AI는 데이터의 의미도 바꾸고 있다. 인터넷의 문서와 이미지가 생성형 AI의 연료였다면, 물리적 AI에는 힘·촉각·관절 상태·거리·속도·공간 관계가 포함된 행동 데이터가 필요하다. 이러한 데이터는 웹에서 손쉽게 수집하기 어렵다. 로봇의 형태와 센서 구성에 따라 데이터 형식도 달라진다. 앞으로의 경쟁력은 모델의 크기만이 아니라 얼마나 다양하고 신뢰할 수 있는 물리적 경험을 확보했는지에 의해 좌우될 가능성이 크다.
공장과 창고에서 시작되는 현실적 확산
물리적 AI가 가장 먼저 확산될 가능성이 큰 장소는 완전히 자유로운 가정이 아니라 공장과 물류창고다. 산업 현장은 작업 목표가 명확하고 공간을 일정 범위 안에서 관리할 수 있으며, 투자 효과도 측정하기 쉽다. 기존 자동화 설비가 반복 작업을 담당했다면, 물리적 AI는 품목과 작업 조건이 자주 바뀌는 영역을 겨냥한다.
제조업에서는 로봇이 서로 다른 크기와 형태의 부품을 분류하고, 조립 과정에서 위치 오차를 스스로 보정하며, 설비 이상을 감지한 뒤 작업 순서를 변경할 수 있다. 다품종 소량생산이 늘어날수록 제품이 바뀔 때마다 로봇을 다시 프로그래밍하는 비용이 커진다. 자연어 지시와 시범 동작만으로 새로운 업무를 학습하는 로봇이 등장하면 생산라인 전환 시간이 줄어든다.
물류 현장에서는 상자의 모양과 무게, 포장 상태가 제각각이라는 문제가 있다. 단단한 상자와 비닐 포장 상품, 깨지기 쉬운 유리 제품을 같은 방식으로 집을 수 없다. 물리적 AI는 시각뿐 아니라 촉각과 힘의 피드백을 이용해 물체의 특성을 추정하고 동작을 조절한다. 주문량의 변화와 작업자의 위치를 고려해 이동 경로를 다시 계산하는 능력도 중요해진다.
건설·농업·에너지 산업에서는 더 복잡한 환경이 기다리고 있다. 건설 현장은 매일 공간 구조가 달라지고, 농장은 날씨와 토양, 작물의 생육 상태가 계속 변한다. 발전소나 해양시설에서는 사람이 접근하기 어려운 위험 구역이 존재한다. 이러한 장소에서 물리적 AI는 사람을 곧바로 대체하는 완전 자율 시스템보다, 위험 작업을 대신하거나 작업자의 판단을 보조하는 형태로 먼저 도입될 가능성이 크다.
물리적 AI의 가치는 인간형 로봇에만 있는 것도 아니다. 사람과 같은 형태는 인간을 위해 설계된 계단·문·도구를 이용하는 데 유리하지만, 특정 작업에는 바퀴형 로봇이나 로봇 팔, 드론, 소프트 로봇이 더 효율적일 수 있다. 창고에서는 무거운 짐을 운반하는 이동형 로봇이, 농장에서는 작물을 손상시키지 않는 유연한 집게가, 재난 현장에서는 좁은 틈을 통과하는 뱀 형태의 로봇이 적합하다. 물리적 AI의 핵심은 인간의 외형을 복제하는 것이 아니라 몸과 과업, 환경의 최적 관계를 찾는 데 있다.
기업의 도입 전략도 달라져야 한다. 범용 로봇 한 대가 모든 일을 해결할 것이라는 기대보다, 작업 범위가 명확하고 데이터가 축적될 수 있는 공정부터 시작해야 한다. 로봇의 성공률만 볼 것이 아니라 작업 중단 시간, 오류 복구 능력, 사람의 개입 횟수, 안전사고 가능성까지 함께 측정해야 한다. 물리적 AI는 소프트웨어 구매가 아니라 설비·공간·업무 절차·인력 운영을 동시에 다시 설계하는 프로젝트이기 때문이다.
의료와 돌봄이 요구하는 더 높은 기준
물리적 AI가 가장 큰 사회적 가치를 만들 수 있는 분야 가운데 하나는 의료와 돌봄이다. 고령화와 의료 인력 부족이 심화되면서 환자 이동, 재활훈련, 물품 운반, 일상생활 보조를 지원할 기술의 필요성이 커지고 있다. 그러나 사람의 몸과 직접 접촉하는 환경은 공장보다 훨씬 높은 수준의 안전성과 섬세함을 요구한다.
돌봄 로봇은 단순히 명령을 정확하게 수행하는 것만으로 충분하지 않다. 사용자의 표정과 자세, 목소리, 움직임의 변화를 감지해야 하며, 갑작스럽게 균형을 잃는 순간에도 적절히 반응해야 한다. 같은 동작이라도 사용자의 신체 상태와 불안감에 따라 속도와 힘을 달리해야 한다. 기술적 성공 여부가 물건을 옮겼는지에 그치지 않고, 사람이 안전함과 존중받고 있음을 느끼는지까지 포함한다.
수술과 재활 분야에서도 물리적 AI는 가능성과 위험을 함께 지닌다. 수술 로봇이 조직의 형태와 움직임을 파악하고 의사의 조작을 보조하면 정밀성을 높일 수 있다. 재활 로봇은 환자의 근력과 반응을 감지해 운동 강도를 조절할 수 있다. 그러나 의료 현장의 데이터는 확보하기 어렵고, 환자마다 해부학적 구조와 상태가 다르다. 드물지만 치명적인 예외 상황을 충분히 학습하기도 쉽지 않다.
따라서 의료 물리적 AI는 단번에 완전 자율화로 이동하기보다 자율성의 수준을 세밀하게 나누는 방식으로 발전해야 한다. 시스템이 스스로 수행할 수 있는 작업, 의료진의 승인이 필요한 작업, 사람이 직접 통제해야 하는 작업을 구분해야 한다. 불확실성이 높아지면 속도를 낮추거나 멈추고 사람에게 판단을 넘기는 능력도 지능의 중요한 일부가 된다.
물리적 AI 시대의 우수한 시스템은 무엇이든 혼자 처리하는 로봇이 아니다. 자신이 아는 것과 모르는 것을 구분하고, 통제권을 언제 인간에게 돌려줘야 하는지 판단하는 로봇이다. 특히 돌봄의 목적은 인간관계를 없애는 데 있지 않다. 반복적이고 육체적으로 부담이 큰 일을 기술이 맡음으로써 돌봄 인력이 대화와 정서적 지원, 전문적 판단에 더 많은 시간을 쓸 수 있게 하는 방향이어야 한다.
정확성보다 무거워지는 안전과 책임
디지털 AI의 오류는 잘못된 정보나 불공정한 추천을 만들 수 있다. 물리적 AI의 오류는 여기에 충돌·낙상·파손·부상이라는 결과를 더한다. 이 때문에 물리적 AI의 안전은 단일한 필터나 사용 규칙으로 해결할 수 없다. 기계적 설계부터 센서, 제어장치, 학습 모델, 운영 절차에 이르는 여러 층의 보호장치가 필요하다.
첫 번째 층은 물리적 안전이다. 로봇의 속도와 힘을 제한하고, 사람이나 장애물을 감지하면 즉시 정지하며, 전력이나 통신이 끊겨도 위험한 동작을 하지 않도록 설계해야 한다. 두 번째 층은 행동 안전이다. AI가 사용자의 지시를 문자 그대로 수행하기 전에 그 행동이 사람과 환경에 위험을 초래하는지 판단해야 한다. 세 번째 층은 운영 안전이다. 누가 로봇을 감독하고, 이상이 발생했을 때 어떤 절차로 개입하며, 사고 기록을 어떻게 보존할지를 정해야 한다.
언어모델에서 나타났던 환각 문제도 행동 영역에서는 훨씬 심각해진다. AI가 존재하지 않는 정보를 만들어내면 답변을 검증할 수 있지만, 존재하지 않는 빈 공간을 있다고 판단하고 로봇 팔을 움직이면 충돌이 발생한다. 자신 있게 틀리는 모델보다 불확실할 때 멈추는 모델이 더 가치 있다. 물리적 AI의 평가 기준에는 평균 성공률뿐 아니라 최악의 상황에서 어떤 행동을 하는지가 포함되어야 한다.
책임 소재도 복잡하다. 로봇이 사고를 일으켰을 때 제조사, AI 모델 개발사, 센서 공급사, 운영 기업, 현장 관리자 가운데 누가 책임을 져야 하는지 명확하지 않을 수 있다. 학습을 통해 행동이 계속 바뀌는 시스템은 출시 당시의 성능만 인증해서도 부족하다. 소프트웨어 업데이트와 데이터 변화가 안전성에 미치는 영향을 지속적으로 점검해야 한다.
사이버보안 역시 물리적 안전과 직결된다. 행동 권한을 가진 AI가 해킹되면 정보 유출을 넘어 설비와 차량, 로봇 자체가 위험 수단으로 바뀔 수 있다. 네트워크에 연결되지 않아도 핵심 기능을 수행할 수 있는 온디바이스 처리, 명령 권한의 분리, 비상 정지 장치, 행동 기록의 위변조 방지가 중요해진다.
사회적 수용성도 기술 성능만으로 결정되지 않는다. 같은 로봇이라도 직원이 위험한 업무를 줄여주는 협력 도구로 받아들이는지, 자신의 행동을 감시하고 일자리를 위협하는 장치로 인식하는지에 따라 도입 결과가 달라진다. 기업은 로봇이 무엇을 할 수 있는지만 설명할 것이 아니라 어떤 데이터를 수집하고, 누가 통제하며, 근로자의 역할이 어떻게 변하는지를 공개해야 한다.
물리적 AI가 다시 묻는 지능의 의미
물리적 AI의 부상은 인공지능 산업의 경쟁 구도를 바꿀 가능성이 크다. 생성형 AI에서는 데이터와 연산능력, 모델 규모가 핵심 자원이었다. 물리적 AI에서는 여기에 센서·반도체·배터리·모터·감속기·소재·로봇 운영 데이터·시뮬레이션 환경이 추가된다. 소프트웨어 기업만으로는 전체 시스템을 완성하기 어렵고, 제조업과 부품기업, 로봇기업, 클라우드기업, 연구기관의 협력이 중요해진다.
특히 제조 기반을 보유한 국가와 기업에는 새로운 기회가 열릴 수 있다. 정밀기계와 전자부품, 자동차, 조선, 배터리, 산업자동화 역량은 물리적 AI의 몸을 구성하는 자산이다. 다만 좋은 하드웨어에 범용 AI를 얹는 것만으로 경쟁력을 확보할 수는 없다. 현장에서 장기간 축적되는 행동 데이터와 작업 지식, 고장과 예외 상황에 대응하는 운영 경험이 기술 격차를 만든다.
고용의 변화도 단순한 대체 논리로만 설명하기 어렵다. 반복적이고 위험한 작업의 일부는 자동화될 수 있지만, 로봇을 가르치고 감독하며 유지하는 역할, 작업환경을 로봇과 인간에게 맞게 재설계하는 역할, 안전성과 윤리를 검증하는 역할이 커진다. 숙련 노동자의 경험은 사라지는 것이 아니라 로봇이 배워야 할 핵심 데이터가 될 수 있다. 문제는 그 지식이 누구의 소유이며, 어떤 보상과 권한 아래 활용되는지를 정하는 데 있다.
물리적 AI는 지능의 정의에도 변화를 요구한다. 시험 문제를 잘 풀고 유창하게 말하는 능력만으로는 현실의 지능을 충분히 설명할 수 없다. 불완전한 감각 정보 속에서 상황을 파악하고, 몸의 한계를 고려해 계획을 세우며, 예상하지 못한 변화에 적응하고, 위험하면 행동을 멈추는 능력이 필요하다. 지능은 아는 것뿐 아니라 적절하게 행동하는 능력이며, 행동의 결과에서 다시 배우는 과정이다.
상용화의 속도를 과장해서도 안 된다. 실험실의 성공적인 시연과 매일 반복되는 현장 운영 사이에는 큰 간격이 존재한다. 몇 분 동안 정교한 동작을 수행하는 것과 수개월 동안 고장 없이 일하는 것은 다른 문제다. 낯선 물체 하나를 다루는 능력과 사람·동물·가구가 끊임없이 움직이는 가정 전체를 이해하는 능력도 같지 않다. 물리적 AI는 빠르게 발전하고 있지만, 범용성과 신뢰성, 비용 효율성을 동시에 확보하기까지는 많은 검증이 필요하다.
그럼에도 방향은 분명하다. AI의 무대가 디지털 공간에서 물리적 세계로 넓어지고 있다. 이 변화의 승자는 가장 사람처럼 생긴 로봇을 먼저 내놓은 기업이 아닐 수 있다. 현실의 복잡성을 이해하고, 몸과 환경의 관계를 학습하며, 실패 가능성을 관리하고, 인간이 신뢰할 수 있는 방식으로 행동하는 시스템을 만든 조직이 앞서게 된다.
생성형 AI가 기계에게 말과 이미지를 만드는 능력을 주었다면, 물리적 AI는 그 지능에 감각과 몸, 행동의 결과를 부여한다. 그 순간 인공지능은 더 이상 화면 속 조언자에 머물지 않는다. 공장의 부품을 집고, 창고의 물건을 옮기며, 농장의 작물을 살피고, 환자의 재활을 돕는 존재가 된다. 기술의 질문도 “AI가 얼마나 똑똑하게 답하는가”에서 “AI가 현실에서 얼마나 안전하고 책임 있게 행동하는가”로 이동한다.
물리적 AI의 미래는 완벽한 자율성만을 향하지 않는다. 인간과 기계가 각자의 강점을 결합해 이전에는 수행하기 어렵거나 위험했던 일을 해결하는 방향으로 나아간다. 행동하는 AI의 시대를 결정할 기준은 화려한 시연이 아니라, 현실의 예외를 견디는 신뢰성과 사람의 통제권을 보존하는 설계가 될 것이다.
Reference
Nature Machine Intelligence, April 2026, Nature Machine Intelligence Editorial Team, From Embodied Intelligence to Physical AI
AI Steps Out of the Screen and Into the Real World
The next frontier of artificial intelligence will not be defined by its ability to produce more natural sentences, but by its capacity to understand complex reality and move safely within it. Physical AI?capable of seeing, hearing, touching, judging and translating those perceptions into real-world action?has begun connecting robotics, manufacturing, logistics, healthcare and caregiving into a single technological movement. The moment AI enters the physical world, its performance standard shifts from providing accurate answers to taking appropriate actions and assuming responsibility for their consequences.
[Key Message]
* AI is expanding from the digital realm into the physical world. Physical AI does more than generate information; it perceives reality, makes decisions and takes direct action.
* Intelligence emerges not only from algorithms but also through interaction among the body and the environment. A robot’s structure, senses and movements are essential components of intelligence, making the body itself part of the computation.
* The competitiveness of physical AI depends on its ability to connect vision, language and action. World models, simulation and data involving touch, force and joint states must be combined so that systems can act appropriately in unfamiliar environments.
* Factories, logistics, healthcare and caregiving are emerging as major application areas for physical AI. The technology is moving beyond the automation of repetitive tasks toward systems that adapt to changing environments and assist human judgment and action.
* The defining standard for AI that acts is not the scale of its intelligence, but its safety and accountability. The ability to stop under uncertainty and return control to a human will determine whether physical AI can earn trust in the real world.
***
From Intelligence on a Screen to Intelligence That Acts
Over the past several years, most advances in artificial intelligence have taken place in the digital realm. Large language models have learned from vast collections of documents to answer questions, while generative AI has produced text, images, music, videos and computer code. Predictive models have calculated weather conditions, financial market movements and the likelihood of disease, while recommendation algorithms have anticipated consumers’ next choices. These systems have demonstrated remarkable capabilities, but their activities have remained largely confined to screens and servers.
The physical world operates under entirely different conditions from the digital realm. If a word is chosen incorrectly in a document, it can be revised. If a robot grasps an object incorrectly, however, the product may be damaged. Awkwardly rendered fingers in a generated image may remain a quality issue, but an incorrect movement by a surgical or caregiving robot could endanger a person. A chatbot can regenerate its answer, whereas an autonomous vehicle facing a pedestrian who suddenly enters the road must select a single course of action within tens of milliseconds.
The editorial “From Embodied Intelligence to Physical AI,” published in *Nature Machine Intelligence* in April 2026, drew attention to this transition. Its central question was clear: what would artificial intelligence need to move beyond predicting, simulating or reasoning about the world and begin acting intelligently within real environments? This was not simply a matter of installing existing AI in a robot. It was a scientific challenge involving how intelligence combines with a body, how that body interacts with its environment, and how the resulting experience reshapes learning and judgment.
“Physical AI” does not refer to one particular algorithm or product. It describes a broad technological system in which sensors that perceive reality, AI models that interpret situations, control systems that plan movement, and robotic bodies that exert physical force operate in a continuous loop. Detecting an object with a camera is not enough. A system must determine what the object is, where it is located, how much force is required to grasp it, and how the surrounding environment will change once it moves. It must then perceive the outcome of its action and immediately revise its plan if reality differs from its prediction.
Until now, artificial intelligence has primarily processed information about the world. It must now act within that world and bear the consequences of its actions. The rise of physical AI represents more than progress in robotics. It marks a transition from AI as a “speaking tool” to AI as an “acting agent.”
Embodied Intelligence: When the Body Becomes Part of Intelligence
Understanding physical AI requires an examination of “embodied intelligence.” Traditional AI research has often treated intelligence as a process of information processing that occurs inside a brain or computer. Under this structure, an internal model calculates data received from the external environment and produces a predefined output. From this perspective, the body is little more than a passive device that carries out commands issued by intelligence.
Embodied intelligence challenges this separation. It views intelligence as something formed not only in the brain but also through bodily structure, sensation, movement and relationships with the environment. The act of picking up a cup offers an intuitive example. A person does not merely calculate the cup’s location. The person simultaneously uses the structure of the finger joints, the friction and pressure felt through the skin, the weight of the cup and the movement of the liquid inside it. If the cup begins to slip, the fingers adjust their grip immediately rather than waiting for a complete visual analysis.
The body is therefore not merely an executor of instructions but part of the computation itself. A bird’s wings and a fish’s fins interact with their environments to produce complex movement naturally. Human muscles and tendons also absorb shocks and maintain balance, reducing the computational burden placed on the brain. The difficulty of the problem an AI must learn changes according to how the robot’s body is designed. Flexible joints and elastic materials can make it easier to absorb collisions and adapt to irregular surfaces. The physical properties of the body assume part of the control algorithm’s role.
The editorial observed that several academic traditions were converging to explain intelligence that acts in the physical world. These included models that internally predict changes in the world, behavior-based robotics emphasizing direct interaction with the environment, ecological psychology connecting perception to possibilities for action, soft robotics using bodily structure and materials themselves, and approaches inspired by biological nervous systems and evolution.
Embodied intelligence and physical AI, however, are not completely interchangeable terms. Embodied intelligence is closer to a scientific and philosophical concept that investigates the role of bodily and environmental interaction in the formation of intelligence. Physical AI broadly encompasses engineering efforts to apply such insights to systems operating in reality. An avatar in a virtual world may also be studied as an embodied agent, but physical AI must contend with actual gravity, friction, collisions, material deformation and sensor errors.
Progress in physical AI will not come from a larger language model alone. The recognition that part of intelligence resides in the body and another part exists in its relationship with the environment is transforming both robot design and AI training.
A New Form of Learning That Connects Perception, Language and Action
Conventional industrial robots have demonstrated exceptional ability in controlled environments. A robotic arm in an automobile factory can quickly and precisely weld components arriving at the same location. If the component’s position changes slightly or an unexpected obstacle appears, however, the robot may have to stop. This approach secures high precision by eliminating as much environmental variability as possible.
The robots envisioned by physical AI are different. Without every possible situation being programmed in advance, they must understand human instructions, adapt to unfamiliar objects and spaces, and respond to changes that occur during an action. When instructed to “clear the dining table,” a robot must distinguish among cups, plates, food waste and fragile objects, decide what to move first, and slow down or stop if a person approaches.
Models that connect vision, language and action are central to this capability. Vision-language-action models do more than understand images and sentences; they generate movements for robots to execute. They interpret their surroundings through camera images, extract goals from natural-language instructions, and determine the direction in which arms and fingers should move or how joints should be controlled. Whereas traditional generative AI predicted the next word, an action model predicts the next movement.
Real-world actions, however, are far more difficult than token generation. Even when handling similar cups, a robot must apply different levels of force depending on whether the cup is made of glass, paper or metal. If it contains liquid, the robot must also regulate its angle and speed. A movement that succeeds on a desk may fail in a moving vehicle or on a swaying vessel. An object that appears clearly visible may be perceived differently because of shadows, reflections, dust or occlusion.
This is where a “world model” becomes important. A world model anticipates the consequences of an action internally before it is performed. It estimates which direction a box will move when pushed, how a door will open when its handle is turned, and whether an object may fall if the robot moves too quickly. If a robot can compare the likely outcomes of multiple actions in advance, it can reduce trial and error and choose a safer plan.
Large-scale simulation and synthetic data have also become essential foundations of physical AI. Teaching a single robot to walk by allowing it to fall thousands of times in reality is time-consuming, expensive and damaging to the equipment. In a virtual environment, by contrast, vast numbers of experiences can be generated quickly by varying gravity, friction, lighting, terrain and object placement. When a factory or warehouse is recreated as a digital twin, different tasks and hazardous scenarios can be tested without interrupting actual operations.
Simulation, however, remains an approximation of reality. A robot that walks perfectly in a virtual environment may fall in the physical world because of slight floor slipperiness, joint wear or sensor delays. Closing this “simulation-to-reality gap” requires virtual training and physical experience to be connected repeatedly. A cyclical structure is needed in which real-world data are reflected in the simulation, policies learned in simulation are tested in physical environments, and the results are used for further revision.
Physical AI is also changing the meaning of data. If online documents and images were the fuel of generative AI, physical AI requires action data incorporating force, touch, joint states, distance, speed and spatial relationships. Such data cannot be collected easily from the web. Their formats also differ according to a robot’s physical design and sensor configuration. Future competitiveness is therefore likely to depend not only on model size but also on the diversity and reliability of the physical experiences a system has acquired.
Practical Adoption Beginning in Factories and Warehouses
The places where physical AI is most likely to spread first are not completely unrestricted homes, but factories and logistics warehouses. Industrial environments have clearly defined objectives, can be managed within controlled boundaries, and allow returns on investment to be measured. Whereas existing automation equipment has handled repetitive operations, physical AI is targeting areas in which products and working conditions change frequently.
In manufacturing, robots could classify components of different shapes and sizes, autonomously correct positional deviations during assembly, detect equipment abnormalities and revise the order of operations. As high-mix, low-volume production expands, the cost of reprogramming robots each time a product changes increases. Robots that learn new tasks from natural-language instructions and demonstrations could significantly reduce production-line conversion time.
In logistics, packages differ widely in shape, weight and condition. A rigid box, a plastic-wrapped product and a fragile glass item cannot be grasped in the same way. Physical AI uses not only vision but also tactile and force feedback to estimate an object’s properties and adjust its movements. The ability to recalculate travel routes by considering changes in order volume and workers’ locations will also become increasingly important.
Even more complex environments await in construction, agriculture and energy. The spatial structure of a construction site changes every day, while farms are continually affected by weather, soil conditions and crop growth. Power plants and offshore facilities contain hazardous areas that are difficult for people to access. In such settings, physical AI is likely to be introduced first not as a fully autonomous system that immediately replaces people, but as a technology that performs dangerous tasks or assists workers’ judgment.
The value of physical AI is not limited to humanoid robots. A humanlike form is advantageous for using stairs, doors and tools designed for people, but wheeled robots, robotic arms, drones or soft robots may be more efficient for particular tasks. Mobile robots may be appropriate for transporting heavy loads in warehouses, flexible grippers for handling crops without damaging them, and snake-shaped robots for passing through narrow gaps at disaster sites. The heart of physical AI lies not in reproducing the human appearance, but in finding the optimal relationship among body, task and environment.
Corporate adoption strategies must also change. Instead of expecting one general-purpose robot to solve every problem, companies should begin with processes that have clearly defined scopes and allow relevant data to accumulate. They must assess not only task success rates but also downtime, error-recovery capabilities, the frequency of human intervention and the possibility of safety incidents. Physical AI is not simply a software purchase. It is a project that redesigns equipment, physical space, work procedures and workforce operations simultaneously.
Higher Standards Required in Healthcare and Caregiving
Healthcare and caregiving are among the fields in which physical AI could create the greatest social value. As populations age and shortages of medical and care workers intensify, the need for technologies that assist with patient transfer, rehabilitation, supply transportation and everyday activities continues to grow. Environments involving direct physical contact with people, however, demand far greater safety and sensitivity than factories.
A caregiving robot must do more than execute instructions accurately. It must detect changes in a user’s facial expression, posture, voice and movement, and respond appropriately if the person suddenly loses balance. Even when performing the same task, the robot must vary its speed and force according to the user’s physical condition and level of anxiety. Technological success cannot be measured solely by whether an object was moved; it must also include whether the person felt safe and respected.
Physical AI presents both possibilities and risks in surgery and rehabilitation. A surgical robot could improve precision by recognizing the shapes and movements of tissues and assisting the surgeon’s control. A rehabilitation robot could detect a patient’s muscular strength and responses and adjust the intensity of exercise accordingly. Yet medical data are difficult to obtain, and every patient’s anatomy and condition differ. It is also difficult to collect enough examples of rare but potentially catastrophic situations.
Medical physical AI should therefore develop by dividing autonomy into carefully calibrated levels rather than moving directly toward complete autonomy. Tasks that a system can perform independently, tasks requiring approval from medical personnel, and tasks that must remain under direct human control should be distinguished. The ability to slow down or stop when uncertainty rises and transfer judgment to a person must become an essential component of intelligence.
A superior system in the age of physical AI will not be a robot that attempts to handle everything alone. It will be a robot that distinguishes what it knows from what it does not know and determines when control should be returned to a human. The purpose of caregiving technology, in particular, should not be to eliminate human relationships. It should allow technology to assume repetitive and physically demanding duties so that care workers can devote more time to conversation, emotional support and professional judgment.
Safety and Responsibility That Carry More Weight Than Accuracy
Errors in digital AI can generate false information or unfair recommendations. Errors in physical AI add collisions, falls, damage and injury to those consequences. For this reason, physical AI safety cannot be secured through a single filter or set of usage rules. Multiple layers of protection are needed, spanning mechanical design, sensors, control devices, learning models and operational procedures.
The first layer is physical safety. A robot’s speed and force must be limited; it must stop immediately when it detects a person or obstacle; and it must be designed to avoid dangerous movements even if power or communication is lost. The second layer is behavioral safety. Before executing a user’s instruction literally, the AI must determine whether the action could endanger people or the environment. The third layer is operational safety. Organizations must define who supervises the robot, what intervention procedures apply when an abnormality occurs, and how incident records will be preserved.
The hallucination problem seen in language models becomes far more serious in the domain of action. If AI produces nonexistent information, its answer can be checked. If it assumes that an occupied space is empty and moves a robotic arm into it, a collision may occur. A model that stops when uncertain is more valuable than one that is confidently wrong. Physical AI evaluations must therefore consider not only average success rates but also how a system behaves under worst-case conditions.
Accountability is also complicated. If a robot causes an accident, responsibility may be difficult to allocate among the robot manufacturer, AI model developer, sensor supplier, operating company and on-site manager. For a system whose behavior continues to change through learning, certification at the time of release is insufficient. The effects of software updates and changing data on safety must be monitored continuously.
Cybersecurity is also directly connected to physical safety. If an AI system with authority to act is compromised, the threat extends beyond information leakage: equipment, vehicles and robots themselves could become instruments of harm. On-device processing that allows essential functions to continue without a network connection, separation of command privileges, emergency-stop mechanisms and tamper-resistant activity records will become increasingly important.
Social acceptance will not be determined by technical performance alone. Adoption outcomes will differ depending on whether employees regard a robot as a collaborative tool that reduces hazardous work or as a device that monitors their behavior and threatens their jobs. Companies must explain not only what a robot can do but also what data it collects, who controls it and how workers’ roles will change.
Physical AI Reopens the Question of What Intelligence Means
The rise of physical AI is likely to transform the competitive landscape of the AI industry. In generative AI, data, computing power and model scale were the primary resources. Physical AI adds sensors, semiconductors, batteries, motors, reduction gears, materials, robotic operational data and simulation environments. Software companies cannot easily complete the entire system on their own, making collaboration among manufacturers, component suppliers, robotics companies, cloud providers and research institutions increasingly important.
This development may create new opportunities for countries and companies with strong manufacturing foundations. Capabilities in precision machinery, electronic components, automobiles, shipbuilding, batteries and industrial automation are assets that form the body of physical AI. Competitive advantage, however, cannot be secured simply by placing a general-purpose AI model on high-quality hardware. The decisive gaps will emerge from action data accumulated over long periods in the field, operational knowledge and experience in responding to failures and exceptional situations.
Changes in employment are also difficult to explain through a simple replacement narrative. Some repetitive and dangerous tasks may be automated, but roles involving the training, supervision and maintenance of robots will expand. So will roles that redesign workplaces for both humans and robots and verify safety and ethical compliance. The experience of skilled workers will not disappear; it may become the essential data that robots must learn. The central issue will be determining who owns that knowledge and under what systems of compensation and authority it is used.
Physical AI also demands a revised definition of intelligence. The ability to solve examination questions and speak fluently is not enough to explain intelligence in reality. A system must interpret situations from incomplete sensory information, develop plans that account for the limitations of its body, adapt to unexpected changes and stop when danger arises. Intelligence is not only the ability to know but also the capacity to act appropriately and learn from the consequences of action.
The pace of commercialization should not be exaggerated. A substantial gap remains between a successful laboratory demonstration and daily operation in the field. Performing an elaborate movement for several minutes is different from working for months without failure. Handling one unfamiliar object is not the same as understanding an entire home in which people, animals and furniture are constantly moving. Physical AI is advancing rapidly, but extensive validation will be required before generality, reliability and cost efficiency can be achieved simultaneously.
The direction, however, is clear. AI’s sphere of activity is expanding from the digital realm into the physical world. The winner of this transformation may not be the company that first introduces the most human-looking robot. Leadership will belong to organizations that understand the complexity of reality, learn the relationship between body and environment, manage the possibility of failure, and create systems that act in ways people can trust.
Generative AI gave machines the ability to create words and images. Physical AI gives that intelligence senses, a body and real consequences. At that point, artificial intelligence no longer remains an adviser confined to a screen. It becomes an entity that grasps components in factories, moves goods through warehouses, monitors crops on farms and assists patients with rehabilitation. The defining technological question therefore shifts from “How intelligently can AI answer?” to “How safely and responsibly can AI act in reality?”
The future of physical AI does not point exclusively toward perfect autonomy. It is moving toward a model in which humans and machines combine their respective strengths to solve tasks that were previously difficult or dangerous. The decisive standards for the age of acting AI will not be spectacular demonstrations, but reliability that withstands real-world exceptions and designs that preserve meaningful human control.
Reference
Nature Machine Intelligence, April 2026, Nature Machine Intelligence, From Embodied Intelligence to Physical AI