🦾
Technology Hub

Robotics & Autonomous Systems

Robotics research spanning manipulation, locomotion, perception, and planning. Humanoid robots, swarm robotics, surgical robots, and autonomous drones.

Engineering / AI
View curated hub

Results for "robotics autonomous manipulation"

227,830 total results — showing 20 from PubMed + NASA ADS + arXiv + OpenAlex
PubMed 2025 Jul

SRT-H: A hierarchical framework for autonomous surgery via language-conditioned imitation learning.

Kim Ji Woong Brian, Chen Juo-Tung, Hansen Pascal, Shi Lucy Xiaoyang, Goldenberg Antony, Schmidgall Samuel, Scheikl Paul Maria, Deguet Anton, White Brandon M, Tsai De Ru, Cha Richard Jaepyeong, Jopling Jeffrey, Finn Chelsea, Krieger Axel

Science robotics

Show Abstract

Research on autonomous surgery has largely focused on simple task automation in controlled environments. However, real-world surgical applications demand dexterous manipulation over extended durations and robust generalization to the inherent variability of human tissue. These challenges remain difficult to address using existing logic-based or conventional end-to-end learning strategies. To address this gap, we propose a hierarchical framework for performing dexterous, long-horizon surgical steps. Our approach uses a high-level policy for task planning and a low-level policy for generating low-level trajectories. The high-level planner plans in language space, generating task-level or corrective instructions that guide the robot through the long-horizon steps and help recover from errors made by the low-level policy. We validated our framework through ex vivo experiments on cholecystectomy, a commonly practiced minimally invasive procedure, and conducted ablation studies to evaluate key components of the system. Our method achieves a 100% success rate across eight different ex vivo gallbladders, operating fully autonomously without human intervention. The hierarchical approach improved the policy's ability to recover from suboptimal states that are inevitable in the highly dynamic environment of realistic surgical applications. This work demonstrates step-level autonomy in a surgical procedure, marking a milestone toward clinical deployment of autonomous surgical systems.

PubMed 2019 Mar

PMK-A Knowledge Processing Framework for Autonomous Robotics Perception and Manipulation.

Diab Mohammed, Akbari Aliakbar, Ud Din Muhayy, Rosell Jan

Sensors (Basel, Switzerland)

Show Abstract

Autonomous indoor service robots are supposed to accomplish tasks, like serve a cup, which involve manipulation actions. Particularly, for complex manipulation tasks which are subject to geometric constraints, spatial information and a rich semantic knowledge about objects, types, and functionality are required, together with the way in which these objects can be manipulated. In this line, this paper presents an ontological-based reasoning framework called Perception and Manipulation Knowledge (PMK) that includes: (1) the modeling of the environment in a standardized way to provide common vocabularies for information exchange in human-robot or robot-robot collaboration, (2) a sensory module to perceive the objects in the environment and assert the ontological knowledge, (3) an evaluation-based analysis of the situation of the objects in the environment, in order to enhance the planning of manipulation tasks. The paper describes the concepts and the implementation of PMK, and presents an example demonstrating the range of information the framework can provide for autonomous robots.

PubMed Review 2022 Sep

Robotic Surgery: A Narrative Review.

Bramhe Sakshi, Pathak Swanand S

Cureus

Show Abstract

In general surgery, the use of robotic and laparoscopic methods has increased. Robotic surgery that requires the least incision has advanced over the years in a short period of time, benefitting both the patient and the surgeon. According to this, robotic platforms and tools are now being used and improved more commonly in general surgery. In a quickly growing and dynamic environment of research and development, the goal of this review is to explore the present and emerging surgical robotic technologies. Future progress in robotics will focus primarily on more durable haptic systems that would provide tactile and kinesthetic input, miniaturisation and micro-robotics, better visual feedback with higher fidelity detail and magnification, and autonomous robots. It is recommended to develop a structured training course with benchmarks for success and evidence-based training strategies. This usually includes a step-by-step progression starting with observation, case aid in programming and manipulation of surgical instruments, learning the basics of robotics in a dry and wet lab setting, attaining non-technical skills on an individual and team level, and monitored modular console training, accompanied by autonomous practice. Prior to independent practice, basic robotics skills and procedural activities must be performed safely and effectively as part of robotic surgical training. It is advised to create a systematic training programme with performance indicators and research-based instructional techniques.

PubMed 2025 Jan

Electromagnetic metamaterial agent.

Hu Shengguo, Li Mingyi, Xu Jiawen, Zhang Hongrui, Zhang Shanghang, Cui Tie Jun, Del Hougne Philipp, Li Lianlin

Light, science & applications

Show Abstract

Metamaterials have revolutionized wave control; in the last two decades, they evolved from passive devices via programmable devices to sensor-endowed self-adaptive devices realizing a user-specified functionality. Although deep-learning techniques play an increasingly important role in metamaterial inverse design, measurement post-processing and end-to-end optimization, their role is ultimately still limited to approximating specific mathematical relations; the metamaterial is still limited to serving as proxy of a human operator, realizing a predefined functionality. Here, we propose and experimentally prototype a paradigm shift toward a metamaterial agent (coined metaAgent) endowed with reasoning and cognitive capabilities enabling the autonomous planning and successful execution of diverse long-horizon tasks, including electromagnetic (EM) field manipulations and interactions with robots and humans. Leveraging recently released foundation models, metaAgent reasons in high-level natural language, acting upon diverse prompts from an evolving complex environment. Specifically, metaAgent's cerebrum performs high-level task planning in natural language via a multi-agent discussion mechanism, where agents are domain experts in sensing, planning, grounding, and coding. In response to live environmental feedback within a real-world setting emulating an ambient-assisted living context (including human requests in natural language), our metaAgent prototype self-organizes a hierarchy of EM manipulation tasks in conjunction with commanding a robot. metaAgent masters foundational EM manipulation skills related to wireless communications and sensing, and it memorizes and learns from past experience based on human feedback.

PubMed 2024 Jan

Finite-time adaptive super-twisting sliding mode control for autonomous robotic manipulators with actuator faults.

Hu Jiabin, Zhang Xue, Zhang Dan, Chen Yun, Ni Hongjie, Liang Huageng

ISA transactions

Show Abstract

This paper proposes a new adaptive super-twisting global integral terminal sliding mode control algorithm for the trajectory tracking of autonomous robotic manipulators with uncertain parameters, unknown disturbances, and actuator faults. Firstly, a novel global integral terminal sliding mode surface is designed to ensure that the tracking errors of autonomous robotic manipulators converge to zero in finite time and the global robustness of the system is also enhanced. Then a new adaptive method is devised to deal with the adverse effect of nonlinear uncertainty. To suppress the chattering phenomenon, the adaptive super-twisting algorithm is used in this paper, which can ensure that the control torque is a continuous input signal. Based on the adaptive mechanism, the adaptive super-twisting global integral terminal sliding mode controller is developed to provide superior control performance. The stability analysis of the system is demonstrated by using the Lyapunov method. Ultimately, the effectiveness of the control scheme is confirmed by a simulation study.

PubMed Review 2024 Mar

Cognitive ergonomics and robotic surgery.

Wong Shing Wai, Crowe Philip

Journal of robotic surgery

Show Abstract

Cognitive ergonomics refer to mental resources and is associated with memory, sensory motor response, and perception. Cognitive workload (CWL) involves use of working memory (mental strain and effort) to complete a task. The three types of cognitive loads have been divided into intrinsic (dependent on complexity and expertise), extraneous (the presentation of tasks) and germane (the learning process) components. The effect of robotic surgery on CWL is complex because the postural, visualisation, and manipulation ergonomic benefits for the surgeon may be offset by the disadvantages associated with team separation and reduced situation awareness. Physical fatigue and workflow disruptions have a negative impact on CWL. Intraoperative CWL can be measured subjectively post hoc with the use of self-reported instruments or objectively with real-time physiological response metrics. Cognitive training can play a crucial role in the process of skill acquisition during the three stages of motor learning: from cognitive to integrative and then to autonomous. Mentorship, technical practice and watching videos are the most common traditional cognitive training methods in surgery. Cognitive training can also occur with computer-based cognitive simulation, mental rehearsal, and cognitive task analysis. Assessment of cognitive skills may offer a more effective way to differentiate robotic expertise level than automated performance (tool-based) metrics.

PubMed 2022 11

Soft robotics for infrastructure protection.

Milana Edoardo

Frontiers in robotics and AI

Show Abstract

The paradigm change introduced by soft robotics is going to dramatically push forward the abilities of autonomous systems in the next future, enabling their applications in extremely challenging scenarios. The ability of soft robots to safely interact and adapt to the surroundings is key to operate in unstructured environments, where the autonomous agent has little or no knowledge about the world around it. A similar context occurs when critical infrastructures face threats or disruptions, for examples due to natural disasters or external attacks (physical or cyber). In this case, autonomous systems may be employed to respond to such emergencies and have to be able to deal with unforeseen physical conditions and uncertainties, where the mechanical interaction with the environment is not only inevitable but also desirable to successfully perform their tasks. In this perspective, I discuss applications of soft robots for the protection of infrastructures, including recent advances in pipelines inspection, rubble search and rescue, and soft aerial manipulation, and promising perspectives on operations in radioactive environments, underwater monitoring and space exploration.

PubMed Review 1990 Sep

Robotic transportation.

Lob W S

Clinical chemistry

Show Abstract

Mobile robots perform fetch-and-carry tasks autonomously. An intelligent, sensor-equipped mobile robot does not require dedicated pathways or extensive facility modification. In the hospital, mobile robots can be used to carry specimens, pharmaceuticals, meals, etc. between supply centers, patient areas, and laboratories. The HelpMate (Transitions Research Corp.) mobile robot was developed specifically for hospital environments. To reach a desired destination, Help-Mate navigates with an on-board computer that continuously polls a suite of sensors, matches the sensor data against a pre-programmed map of the environment, and issues drive commands and path corrections. A sender operates the robot with a user-friendly menu that prompts for payload insertion and desired destination(s). Upon arrival at its selected destination, the robot prompts the recipient for a security code or physical key and awaits acknowledgement of payload removal. In the future, the integration of HelpMate with robot manipulators, test equipment, and central institutional information systems will open new applications in more localized areas and should help overcome difficulties in filling transport staff positions.

NASA ADS 2024-05-00
2 citations

Learning strategies for underwater robot autonomous manipulation control

Huang, Hai, Jiang, Tao, Zhang, Zongyu, Sun, Yize, Qin, Hongde, Li, Xinyang, Yang, Xu

Journal of the Franklin Institute

Show Abstract

Autonomous manipulation operations represent the high intelligent coordination from robotic vision and control, it is also a symbol of the advances of robotic intelligence. The limitations of visual sensing and the increasingly complex experimental conditions make autonomous manipulation operations more difficult, particularly for deep reinforcement learning methods, which can enhance robotic control intelligence but require a lot of training process. Due to the high-dimensional continuous state space and continuous action space characteristics of underwater operations, this paper adopts a policy-based reinforcement learning method as the foundational approach. To address the issues of instability and low convergence efficiency in traditional policy-based reinforcement learning algorithms during the learning process, this paper proposes a novel policy learning method. This method adopts the Proximal Policy Optimization algorithm (PPOClip) and optimizes it through an actor-critic network. The aim is to improve the stability and effectiveness of convergence in the learning process. In the underwater training environment, a new reward shaping scheme has been designed to address the issue of reward sparsity during the training process. The manually crafted dense reward function is utilized as attractive and repulsive potential functions for goal manipulation and obstacle avoidance. On the highly complex underwater manipulation and training environment, transferred learning algorithm has been established to reduce the training times and compensate the differences between the simulation and experiment. Simulations and tank experiments have verified the performance of the proposed strategy learning method.

NASA ADS 2024-10-00
75 citations

Generalizable Humanoid Manipulation with 3D Diffusion Policies

Ze, Yanjie, Chen, Zixuan, Wang, Wenhao, Chen, Tianyi, He, Xialin, Yuan, Ying, Peng, Xue Bin, Wu, Jiajun

arXiv e-prints

Show Abstract

Humanoid robots capable of autonomous operation in diverse environments have long been a goal for roboticists. However, autonomous manipulation by humanoid robots has largely been restricted to one specific scene, primarily due to the difficulty of acquiring generalizable skills and the expensiveness of in-the-wild humanoid robot data. In this work, we build a real-world robotic system to address this challenging problem. Our system is mainly an integration of 1) a whole-upper-body robotic teleoperation system to acquire human-like robot data, 2) a 25-DoF humanoid robot platform with a height-adjustable cart and a 3D LiDAR sensor, and 3) an improved 3D Diffusion Policy learning algorithm for humanoid robots to learn from noisy human data. We run more than 2000 episodes of policy rollouts on the real robot for rigorous policy evaluation. Empowered by this system, we show that using only data collected in one single scene and with only onboard computing, a full-sized humanoid robot can autonomously perform skills in diverse real-world scenarios. Videos are available at https://humanoid-manipulation.github.io .

NASA ADS 2023-00-00
2 citations

Six-Dimensional Target Pose Estimation for Robot Autonomous Manipulation: Methodology and Verification

Wang, Rui, Su, Congjia, Yu, Hao, Wang, Shuo

IEEE Transactions on Cognitive and Developmental Systems

Show Abstract

The autonomous and precise grasping operation of robots is considered challenging in situations where there are different objects with different shapes and postures. In this study, we proposed a method of 6-D target pose estimation for robot autonomous manipulation. The proposed method is based on: 1) a fully convolutional neural network for scene semantic segmentation and 2) fast global registration to achieve target pose estimation. To verify the validity of the proposed algorithm, we built a robot grasping operation system and used the point cloud model of the target object and its pose estimation results to generate the robot grasping posture control strategy. Experimental results showed that the proposed method can achieve a six-degree-of-freedom pose estimation for arbitrarily placed target objects and complete the autonomous grasping of the target. Comparative experiments demonstrated that the proposed target pose estimation method achieved a significant improvement in average accuracy and real-time performance compared with traditional methods.

NASA ADS 1989-06-00
375 citations

A new technique for fully autonomous and efficient 3D robotics hand/eye calibration

Tsai, R. Y., Lenz, R. K.

IEEE Transactions on Robotics Automation

Show Abstract

The authors describe a novel technique for computing position and orientation of a camera relative to the last joint of a robot manipulator in an eye-on-hand configuration. It takes only about 100+64N arithmetic operations to compute the hand/eye relationship after the robot finishes the movement, and incurs only additional 64 arithmetic operations for each additional station. The robot makes a series of automatically planned movements with a camera rigidly mounted at the gripper. At the end of each move, it takes a total of 90 ms to grab an image, extract image feature coordinates, and perform camera extrinsic calibration. After the robot finishes all the movements, it takes only a few milliseconds to do the calibration. A series of generic geometric properties or lemmas are presented, leading to the derivation of the final algorithms, which are aimed at simplicity, efficiency, and accuracy while giving ample geometric and algebraic insights. Critical factors influencing the accuracy are analyzed, and procedures for improving accuracy are introduced. Test results of both simulation and real experiments on an IBM Cartesian robot are reported and analyzed.

arXiv 2020-10-13

Real-Time Deep Learning Approach to Visual Servo Control and Grasp Detection for Autonomous Robotic Manipulation

Eduardo Godinho Ribeiro, Raul de Queiroz Mendes, Valdir Grassi

Journal: Robotics and Autonomous Systems, publisher: Elsevier, volume number: 139, year: 2021, page number: 103757

Show Abstract

In order to explore robotic grasping in unstructured and dynamic environments, this work addresses the visual perception phase involved in the task. This phase involves the processing of visual data to obtain the location of the object to be grasped, its pose and the points at which the robot`s grippers must make contact to ensure a stable grasp. For this, the Cornell Grasping dataset is used to train a convolutional neural network that, having an image of the robot`s workspace, with a certain object, is able to predict a grasp rectangle that symbolizes the position, orientation and opening of the robot`s grippers before its closing. In addition to this network, which runs in real-time, another one is designed to deal with situations in which the object moves in the environment. Therefore, the second network is trained to perform a visual servo control, ensuring that the object remains in the robot`s field of view. This network predicts the proportional values of the linear and angular velocities that the camera must have so that the object is always in the image processed by the grasp network. The dataset used for training was automatically generated by a Kinova Gen3 manipulator. The robot is also used to evaluate the applicability in real-time and obtain practical results from the designed algorithms. Moreover, the offline results obtained through validation sets are also analyzed and discussed regarding their efficiency and processing speed. The developed controller was able to achieve a millimeter accuracy in the final position considering a target object seen for the first time. To the best of our knowledge, we have not found in the literature other works that achieve such precision with a controller learned from scratch. Thus, this work presents a new system for autonomous robotic manipulation with high processing speed and the ability to generalize to several different objects.

arXiv 2021-12-03

A Survey of Robot Manipulation in Contact

Markku Suomalainen, Yiannis Karayiannidis, Ville Kyrki

Robotics and Autonomous Systems, Volume 156, 2022, 104224, ISSN 0921-8890,

Show Abstract

In this survey, we present the current status on robots performing manipulation tasks that require varying contact with the environment, such that the robot must either implicitly or explicitly control the contact force with the environment to complete the task. Robots can perform more and more manipulation tasks that are still done by humans, and there is a growing number of publications on the topics of 1) performing tasks that always require contact and 2) mitigating uncertainty by leveraging the environment in tasks that, under perfect information, could be performed without contact. The recent trends have seen robots perform tasks earlier left for humans, such as massage, and in the classical tasks, such as peg-in-hole, there is a more efficient generalization to other similar tasks, better error tolerance, and faster planning or learning of the tasks. Thus, in this survey we cover the current stage of robots performing such tasks, starting from surveying all the different in-contact tasks robots can perform, observing how these tasks are controlled and represented, and finally presenting the learning and planning of the skills required to complete these tasks.

arXiv 2026-07-16

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories

Xiaomi Robotics Team, Jun Guo, Piaopiao Jin, Jason Li, Peiyan Li, Yingyan Li, Futeng Liu, Wanli Peng, Optimus Qin, Yifei Su, Nan Sun, Qiao Sun, Runze Suo, Heyun Wang, Yunhong Wang, Rujie Wu, Caoyu Xia, Lina Zhang, Jack Zhao, Guoliang Chen, Wenlong Chen, Xinze He, Bin Li, Qing Li, Zhuorong Li, Heng Qu, Wenxuan Song, Diyun Xiang, Yifan Xie, Peiran Xu, Hangjun Ye, Wen Ye, Han Zhao, Quanyun Zhou

arXiv:2607.15330v2 [cs.RO]

Show Abstract

We present Xiaomi-Robotics-1, a foundational vision-language-action (VLA) model capable of (1) following diverse language instructions to perform a wide range of mobile manipulation tasks in unseen environments out-of-the-box, and (2) efficiently adapting to novel downstream tasks with minimal fine-tuning data. We propose a two-stage training recipe consisting of pre-training and post-training. During pre-training, we imbue the model with broad and generalizable action-generation capabilities by training on over 100k hours of real-world manipulation trajectories collected via UMI devices. Crucially, we develop a scalable auto-labeling pipeline that annotates trajectory clips with natural languages describing scene state transitions, providing rich and precise conditioning for action learning. During post-training, we aim to align these capabilities with robot embodiments and imperative instructions that humans naturally use to prompt robots. Extensive experiments demonstrate strong scaling behavior. Xiaomi-Robotics-1 consistently improves with increased data scales and model sizes during pre-training. This scaling behavior directly transfers to post-training, where a stronger pre-training model yields better out-of-the-box real-robot performance in unseen environments. Furthermore, Xiaomi-Robotics-1 serves as a strong robot foundation policy that can be efficiently fine-tuned on complex, dexterous tasks with high data efficiency. Across multiple simulation benchmarks, Xiaomi-Robotics-1 outperforms state-of-the-art methods. Notably, it establishes a new state-of-the-art with a 57.4% success rate on RoboCasa365, surpassing the previous best of 46.6%. Furthermore, it achieves an average score of 20.07 on RoboDojo, significantly outperforming the prior state-of-the-art (13.07). Code and model checkpoints will be released. Project page: https://robotics.xiaomi.com/xiaomi-robotics-1.html

arXiv 2026-01-02

Vision-based Goal-Reaching Control for Mobile Robots Using a Hierarchical Learning Framework

Mehdi Heydari Shahna, Pauli Mustalahti, Jouni Mattila

Robotics and Autonomous Systems 206 (2026) 105710

Show Abstract

Reinforcement learning (RL) has strong potential in robotics, but exploration-based training complicates safe deployment on large-scale robots. For such applications, this paper proposes a novel hierarchical goal-reaching framework that integrates stereo visual pose estimation, constrained RL-based motion planning, actuator-level robust adaptive control (RAC), and supervisory safe-return logic. Stereo visual localization is used as the real-time pose-estimation interface with loop closing, map fusion, and relocalization. The RL planner generates smooth, feasible goal-reaching references using a problem-specific reward structure and motion constraints that promote goal progress, reduce oscillations, preserve vision-consistent smoothness, and respect the mechanical limits of a heavy skid-steered robot. At the actuation layer, a scaled conjugate-gradient (SCG)-trained deep neural network (DNN) approximates a quasi-static actuator feedforward map from wheel-speed data to nominal control input. This feedforward map is combined with a logarithmic-barrier-based RAC to compensate for residual modeling errors, slip-induced disturbances, and bounded mismatch between the nominal map and real actuator response. For the actuator-level wheel-tracking subsystem, uniformly ultimately bounded tracking with exponential convergence to a disturbance-dependent residual set is established under bounded uncertainty. A logarithmic safety supervisor monitors execution, detects unsafe operating conditions, including faults and localization inconsistencies, and switches the robot to safe-return mode. Experiments on a 6000 kg robot over asphalt and loose-soil terrain demonstrate approximately 3--4 cm final-position root mean square error (RMSE), accurate tracking of RL-generated commands, improved actuator-level performance over two RAC baselines, and successful autonomous recovery after fault injection.

OpenAlex 2012-10-01
120 citations

An integrated system for autonomous robotics manipulation

J. Andrew Bagnell, Felipe Lira de Sá Cavalcanti, Lei Cui, Thomas Galluzzo, Martial Hebert, Moslem Kazemi, Matthew Klingensmith, Jacqueline Libby, Tian Yu Liu, Nancy S. Pollard, Mihail Pivtoraiko, Jean‐Sebastien Valois, Ranqi Zhu

Show Abstract

We describe the software components of a robotics system designed to autonomously grasp objects and perform dexterous manipulation tasks with only high-level supervision. The system is centered on the tight integration of several core functionalities, including perception, planning and control, with the logical structuring of tasks driven by a Behavior Tree architecture. The advantage of the implementation is to reduce the execution time while integrating advanced algorithms for autonomous manipulation. We describe our approach to 3-D perception, real-time planning, force compliant motions, and audio processing. Performance results for object grasping and complex manipulation tasks of in-house tests and of an independent evaluation team are presented.

OpenAlex 2017-05-01
1469 citations

Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates

Shixiang Gu, Ethan Holly, Timothy Lillicrap, Sergey Levine

Show Abstract

Reinforcement learning holds the promise of enabling autonomous robots to learn large repertoires of behavioral skills with minimal human intervention. However, robotic applications of reinforcement learning often compromise the autonomy of the learning process in favor of achieving training times that are practical for real physical systems. This typically involves introducing hand-engineered policy representations and human-supplied demonstrations. Deep reinforcement learning alleviates this limitation by training general-purpose neural network policies, but applications of direct deep reinforcement learning algorithms have so far been restricted to simulated settings and relatively simple tasks, due to their apparent high sample complexity. In this paper, we demonstrate that a recent deep reinforcement learning algorithm based on off-policy training of deep Q-functions can scale to complex 3D manipulation tasks and can learn deep neural network policies efficiently enough to train on real physical robots. We demonstrate that the training times can be further reduced by parallelizing the algorithm across multiple robots which pool their policy updates asynchronously. Our experimental evaluation shows that our method can learn a variety of 3D manipulation skills in simulation and a complex door opening skill on real robots without any prior demonstrations or manually designed representations.

OpenAlex 2018-11-01
86 citations

Visual Manipulation Relationship Network for Autonomous Robotics

Hanbo Zhang, Xuguang Lan, Xinwen Zhou, Zhiqiang Tian, Yang Zhang, Nanning Zheng

Show Abstract

Robotic grasping is one of the most important fields in robotics, in which great progress has been made in recent years with the help of convolutional neural network (CNN). However, including multiple objects in one scene can invalidate the existing CNN-based grasp detection algorithms, because manipulation relationships among objects are not considered, which are required to guide the robot to grasp things in the right order. This paper presents a new CNN architecture called Visual Manipulation Relationship Network (VMRN) to help robots detect targets and predict the manipulation relationships in real time, which ensures that the robot can complete tasks in a safe and reliable way. To implement end-to-end training and meet real-time requirements in robot tasks, we propose the Object Pairing Pooling Layer (OP2L) to help to predict all manipulation relationships in one forward process. Moreover, in order to train VMRN, we collect a dataset named Visual Manipulation Relationship Dataset (VMRD) consisting of 5185 images with more than 17000 object instances and the manipulation relationships between all possible pairs of objects in every image, which is labeled by the manipulation relationship tree. The experimental results show that the new network architecture can detect objects and predict manipulation relationships simultaneously and meet the real-time requirements in robot tasks.