Enter your keyword

Autonomous Tram System: An Innovative Collaboration between ITB and INKA

Autonomous Tram System: An Innovative Collaboration between ITB and INKA

Bambang Riyanto Trilaksono
School of Electrical Engineering and Informatics-ITB, Center for Artificial Intelligence-ITB

Ari Wibowo
School of Electrical Engineering and Informatics-ITB

A. Pendahuluan
In an era of rapid technological development, transportation is one sector that continues to undergo transformation. In an effort to provide innovative and environmentally friendly solutions, the Bandung Institute of Technology (ITB) and PT Industri Kereta Api (INKA) have partnered to develop an autonomous tram system. This initiative not only creates a modern transportation solution but also drives technological progress in Indonesia. Safe and comfortable public transportation is a crucial element of urban infrastructure. Considering several factors, such as traffic congestion, environmental concerns, and rapid urban growth, public transportation must be the most attractive way to get around. The growing interest in electric mobility and greater sensitivity to environmental issues have supported the spread of trams in recent years (the City Government of Bogor, the Province of Bali, and the City of Surabaya) have expressed interest in using trams as their urban transportation (source: PT INKA internal). This is likely to be followed by other cities in Indonesia, thus increasing the potential for tram production. The term tram is used for public transportation trains that run parallel to and on the road. This means that trams must blend in with other means of transportation, such as cars and motorcycles, as well as pedestrians. During operation, trams must limit their speed to avoid accidents with other modes of transportation using the road. In addition to speed limits, trams are also typically limited to one or two cars to avoid disrupting traffic flow. They generally have a capacity of between 125 and 250 passengers.

Currently, train accidents in Indonesia are high, mainly due to driver fatigue or drowsiness. Proposed solutions include adding driver assistance systems to trains or trams, or even developing autonomous trams to improve the safety and efficiency of public transportation. Autonomous vehicles are considered an effective solution to address the problem of accidents caused by suboptimal driver conditions. This article will discuss research and innovation aspects related to autonomous vehicles, with a focus on the development of autonomous trams. The discussion includes the development of autonomous systems, including deep learning and the software and hardware designed. Test results of the assisted driving system on autonomous trams will also be described, followed by conclusions and lessons learned from the development of autonomous trams with artificial intelligence.

B. Trem Otonom
Based on the above considerations, the Bandung Institute of Technology (ITB) has established a strategic partnership with PT Industri Kereta Api (PT INKA) and an artificial intelligence research company in an innovative research project entitled "Development of an Autonomous System Using Artificial Intelligence for Autonomous Trams." This project received funding from the Indonesian Endowment Fund for Education (LPDP) under the productive innovative research scheme. Within the context of this partnership, ITB takes a leading role in the development of the autonomous system. This includes the design and development of the AI pipeline, software, and hardware, including sensors with embedded graphical processing units (GPUs). In addition, PT INKA contributes by providing electric trams and drive-by-wire systems, as well as providing a testing environment in Madiun. Furthermore, artificial intelligence research is involved in the development of specialized AI models for facial recognition and driver attention warnings. They also play a crucial role in monitoring and analyzing data used in system engineering. A research and innovation roadmap has been designed to encompass the research stages that have been conducted, are ongoing, and will be conducted. Figure 1 visually illustrates this roadmap.

Figure 1. Autonomous Tram Development Research Roadmap (2019–2027)

In the 2021-2023 period, the main focus of this project is the development of an autonomous tram system. After that, in the 2024-2027 period, this project will expand its scope to include the development of Autonomous Rail Rapid Transit (ART). Thus, this project not only provides a solution for autonomous trams, but also plans to implement autonomous technology in a wider range of rail transportation modes. The overall project has a long-term goal of improving the safety and efficiency of public transportation in Indonesia. As shown in Figure 1, several studies have been conducted related to the detection and classification and tracking of traffic objects based on cameras as an important element of autonomous vehicles [1][2][18][19][20]. Several deep learning models have been developed to support this research, one of which is based on Yolo to detect various traffic objects, as is often found in Indonesia. Deep learning is an important paradigm in artificial intelligence that has developed in recent years and is based on the model and working principles of the brain of living beings in general, and humans in particular, which is organized into a large number of layers of neurons [2][3][4][9].

C. Autonomous Tram System
This autonomous tram system is supported by advanced technologies such as smart sensors, artificial intelligence (AI), and real-time data processing. Smart sensors installed on the tram can detect road conditions, surrounding vehicles, and changes in the surrounding environment. Information obtained from these sensors will be processed by the artificial intelligence system to make appropriate decisions, including navigation and speed control of the tram. As an autonomous vehicle, the autonomous tram must be able to detect objects based on the sensors used and be able to avoid obstacles. This is important because the autonomous tram operates in a mixed road environment with various types of vehicles such as cars, public transportation, motorcycles, bicycles, buses, trucks, and pedestrians. Various sensors are designed and used on the autonomous tram, including camera sensors, LIDAR sensors, and radar sensors. Each sensor has unique characteristics that complement each other. The autonomous tram must also be able to operate within the permitted speed limits on certain rail sections. This is necessary to prevent the tram from becoming dangerous, especially in situations such as steep slopes. The autonomous tram development process is carried out in two stages: the design of tram driving assistance in the 2021-2022 period and the development of autonomous trams in the 2022-2023 period.

Tram driving assistance can be considered a semi-autonomous system that functions as an aid to the driver, especially in certain conditions that are considered dangerous. This system needs to be developed, especially for trams operating in Indonesia, considering the tendency of some traffic users in this country who do not always obey traffic regulations. The autonomous system in tram driving assistance needs to be equipped with the ability to recognize the driver's face and provide a driver's attention warning. This capability is crucial in situations where the driver experiences fatigue or loses concentration while operating the tram. The attention warning system must be able to detect the driver's lack of focus. In addition, during the development stage of the autonomous tram, several features have been added to the autonomous system. These features include adaptive cruise control, obstacle avoidance, and autonomous emergency braking. All of these features are designed to improve the safety and efficiency of the autonomous tram, thus making a positive contribution to public transportation in Indonesia.

To realize tram driving assistance, a deep learning model was developed that has the ability to detect objects around the tram and track them using camera sensors, LiDAR, and radar implemented in the perception module of the developed autonomous system. The types of objects that need to be detected are motorcycles, cars, public transportation, trucks, buses, and pedestrians. A deep learning model based on convolutional neural networks was also developed to identify faces through a camera installed in the driver's compartment and can identify who is the driver. A warning system was also realized using a deep learning model to detect whether the driver is drowsy or distracted.

Figure 2. Block Diagram of Autonomous Tram System Design

The block diagram of the autonomous system design for the developed autonomous tram is shown in Figure 2. As shown in the figure, the autonomous system for the autonomous tram is equipped with several sensors, such as cameras, LiDAR, and radar. These types of sensors are necessary because the autonomous tram operates in a mixed traffic environment, or in other words, operates alongside or in the middle of the highway together with various other types of vehicles, such as bicycles, motorcycles, public transportation, cars, buses, and trucks, or pedestrians. The installation of sensors on the autonomous tram and their placement are shown in Figure 3 (a), while Figure 3 (b) is the placement of sensors on a testbed vehicle for simulation and testing before being installed on the tram. The use of camera, LiDAR, and radar sensors in traffic conditions such as these is necessary to perceive the environment around the autonomous tram. Specifically, the perception system equipped with these sensors has the ability to detect and classify surrounding objects. Object detection is carried out by forming a two- or three-dimensional bounding box using deep learning based on camera, LiDAR, and radar sensors. An example of the inference of the developed deep learning model on video captured by the camera is shown in Figure 4.

Camera-based object detection methods work by using images as input to generate information about objects in the autonomous vehicle environment. The development of this object detection method is very important in autonomous vehicles, because it requires detection with high accuracy and speed. Camera-based object detection methods can be divided into two types, namely single-image or monocular camera-based methods (monocular-camerabased) and multiple-image or stereo camera-based methods (stereo-camerabased). Single-image based methods only use one image to obtain information about 2D objects, while stereo methods use two or more images to generate more accurate information and can even produce 3D object detection. Some advantages of monocular camera-based object detection methods include ease of implementation, lower costs compared to stereo camera sensors, LiDAR or radar. However, some disadvantages of this method are that it cannot obtain object depth well, limited accuracy especially at long distances, is affected by light and weather conditions, and is susceptible to occlusion.

Gambar 3. Gambar sebelah kiri penempatan sensor pada trem, sebelah kanan penempatan sensor pada kendaraan testbed
Gambar 4. Deteksi objek pada video yang ditangkap oleh kamera dengan menggunakan deep learning, gambar sebelah kiri mengenali rambu perkeretaapian, gambar sebelah kanan mengenali objek-objek yang ada di sekitar trem

The camera used in this study is the Sekonix SF3325-100 camera, the camera is a type of wide-angle monocular camera that has a horizontal viewing angle of 60°. This camera is capable of processing images at speeds of up to 30 images per second or frames per second (FPS). In addition, the Sekonix SF3325-100 camera is a type of camera that is capable of working at storage temperatures up to 105°C and protection resistance is IP69K. Meanwhile, for the LiDAR used the production type from Velodyne Inc with the type VLP-32C. This LiDAR is capable of working at a speed of 5 – 20 Hz with a number of channels of 32. This LiDAR produces scans with a horizontal viewing angle of 360°. The next sensor used for the perception system is radar, the type of radar used is the Continental ARS430RDI radar. This type of radar is an automotive radar that operates at a frequency of 77 GHz. The ARS430RDI has 2 sensing areas that work simultaneously. The ARS430RDI can detect objects at elevations of 14 and 20 degrees from the ground surface for the far and near sensing regions, respectively. The list of sensors used in the autonomous tram perception system is shown in Figure 5.

Figure 5. Camera, radar, and lidar sensors used in the autonomous tram perception system.

In the perception system, image segmentation has also been developed, which functions to increase the "intelligence" of the perception system. For example, by using one of the methods in image segmentation, namely deep learning-based panoptic image segmentation, the autonomous system can distinguish images that are trees, buildings, air, vehicle objects around the autonomous tram, and road areas (drivable areas). The results of the inference of the deep learning model for this image segmentation are shown in Figure 6. The left image shows the image segmentation of vehicle objects and trees. The right image shows the image segmentation to determine the area/region where the vehicle can operate. As seen in this image, the developed deep learning model can build image separation of vehicle objects, trees, and road areas that can be used by the autonomous vehicle. In addition to the ability to detect objects, a traffic sign recognition system is also developed based on images captured by the camera. This system also uses a deep learning model, so it has the ability to detect signs or slogans used along the autonomous railway, including the ability to detect signs in rather extreme situations, such as dim lighting, for example in the afternoon, or in situations when there is light rain.

This perception system is further linked to the localization and mapping systems. The localization system is needed to determine the location coordinates of the autonomous tram on the map. Accurate location determination is crucial for autonomous vehicles. This system calculates location coordinates based on GPS and IMU sensors with a certain level of accuracy, namely in the case of the autonomous tram being developed, the accuracy is below 10 cm. There are two types of GPS used: the first utilizes satellites, and the second is RTK-based. For maps, the open-source Open Streetmap is used. The autonomous tram's location information is then shown on the map and displayed on a display.

Based on object detection, localization, and mapping systems, the autonomous system's decision-making system will make several decisions, such as accelerating and decelerating the tram, sounding the buzzer in the driver's compartment, or sounding the horn. The decision-making module consists of a safety assessment, decision-making, and control system. Safety assessment allows the autonomous system to assess the level of danger of a particular situation around the autonomous tram. This assessment is expressed in the form of a probability value of the danger of a particular situation, based on the predicted trajectory of the object (vehicle) and the position and speed of the autonomous tram. This system is designed to realize important features of autonomous vehicles, such as emergency braking and collision avoidance. An adaptive cruise control feature has also been developed that allows the autonomous tram to set a safe distance (and speed) from other trams traveling in front of it. One of the control methods used is MPC (Model Predictive Control).

Figure 6. Segmentation inference on images using deep learning.

Some advantages of monocular camera-based object detection methods include ease of implementation, lower cost compared to stereo camera sensors, LiDAR, or radar. However, some disadvantages of this method are that it cannot obtain object depth well, limited accuracy especially at long distances, is affected by light and weather conditions, and is susceptible to occlusion. Therefore, LiDAR and radar sensors are needed to obtain more complete 3D object information. Object detection based on LiDAR alone obtains more information than using only a camera. By calculating laser return patterns, LiDAR sensors can create a 3D point cloud representation of the surrounding environment. This LiDAR-based object detection system can generate 3D bounding boxes that show the position, size, and orientation of objects around an autonomous vehicle (Simony, M. et al., 2018). LiDAR-based object detection methods have several advantages compared to camera-based methods, including high accuracy and reliability, independence from lighting conditions, and the ability to detect objects at longer distances. This method is also less affected by object occlusion and can detect objects with low reflectivity or dark-colored vehicles [22].

This research also develops a model that utilizes camera fusion with LiDAR, 3D object detection based on camera fusion and LiDAR is very popular recently for the development of autonomous vehicles. This method is also known as the use of multi-sensors. The use of methods that only rely on one type of sensor such as LiDAR or camera often results in less than satisfactory detection in some cases, especially when lighting conditions are poor or there are obstructing objects. Therefore, the use of multi-sensors is a better solution to obtain more accurate and reliable results. The fusion of both data information from both camera and LiDAR sensors is used to obtain more complete and detailed 3D information about the surrounding environment. This approach allows to combine the advantages of each sensor. LiDAR can provide very accurate information about the distance and position of objects around the autonomous vehicle, which is able to overcome the problem of limited range and the inability to identify objects in low lighting conditions on the camera, while the camera is able to provide richer information about visual characteristics such as color, texture, and pattern which are limitations of LiDAR data [23].

Based on the specifications of the sensor hardware used and the placement of the sensors on the tram, this research has limitations in the field of view (FOV) of the system being developed. Simply put, the FOV from the output of this research perception system can be depicted in Figure 7 below.

Gambar 7. Field of View (a) Sensor LiDAR (b) Sensor Kamera

The 3D detection system of LiDAR will generate a 3D bounding box (x,y,z,w,h,l) based on its classification (k) along with orientation (𝜃) and detection confidence score (s), as shown in Figure 8. The 3D bounding box in LiDAR coordinates can be accurately projected onto the image plane by using calibration parameters between the camera and LiDAR, as shown in Figure 9.

Gambar 8. Contoh hasil deteksi objek 3D menggunakan LiDAR

Furthermore, in camera-based object detection, 2D object detection is faster than 3D object detection. In this study, camera-LiDAR fusion was performed at the late-fusion or decision-level stage. The fusion method at this stage focuses on improving average precision (AP) by improving or correcting each individual detector. Therefore, camera detection eliminates the need for 3D object detection.

Figure 9. Illustration of the projection of 3D object detection results onto the image plane.

However, since the output data of 2D and 3D object detection are different, the association process will be carried out on the same plane, namely the image plane. This can be done because the projection of the 3D bounding box onto the 2D image plane can be done very precisely with the help of several calibration parameters [24].

D. Implementation of Autonomous Systems
The entire AI pipeline and algorithms used in this designed autonomous system are implemented on Nvidia GPU Driveworks (Pegasus). Nvidia GPU Driveworks, also known as Pegasus, is an embedded GPU developed for autonomous vehicles, containing multiple GPUs on a single board. It is important to note that in the implementation and deployment of deep learning models, simply developing them as is commonly done in academic environments is not sufficient. These developed AI models need to be deployed and tested on the computing platform used for the system deployment, in this case, the Embedded GPU Driveworks. In general, the development of deep learning models in this research and innovation project is carried out through several stages:

  1. Deep learning models are developed through transfer learning from pretrained models or modifications of existing deep learning models developed by researchers. These models are trained using online datasets such as Kitti and NuScene. They are then trained on local datasets collected and developed by the research team. These deep learning models are developed on a server with multiple GPUs. Development is performed using TensorFlow or PyTorch.
  2. Deep learning models were developed on servers and GPUs using the same software environment (Nvidia Driveworks) as the Nvidia Embedded GPU Drive (Pegasus). In this case, the same version of TensorRT was used.
  3. Implementation of AI pipeline and related algorithms on Nvidia Embedded GPU Driveworks using TensorRT.

Specifically, in Step 3), an assessment is carried out not only regarding the performance of the developed deep learning model, for example, accuracy, recall, f1-score, and others, but also the computational speed, for example in the form of fps (frames per second). This evaluation is very important because the deep learning model implemented on the Embedded GPU will be applied to a safety-critical system, namely an autonomous tram, where the response time of various models or algorithms is very important to pay attention to. In this case, there is often a compromise between the performance of the deep learning model and its computational time. In some cases, we are often unable to develop and use the most sophisticated and latest deep learning models due to limitations of the TensorRT version used on the Embedded GPU. In some cases, a compromise or trade-off in AI design and deployment is necessary.

In general, the development of deep learning and AI pipeline models and their simulations were carried out at ITB and Artificial Intelligence Research. Meanwhile, integration on the tram platform was carried out at PT INKA, Madiun. Environmental data collection using camera, lidar, and radar sensors was carried out in several locations: 1) Bandung Streets, 2) Solo (on the Bathara Kresna tram rail section), 3) PT INKA. Street data collection in Bandung was intended to obtain video and lidar data for object detection, including in rather extreme conditions, namely in rainy and rather dark conditions (afternoon approaching night). Data collection in Solo used a dressin similar to an electric tram that had its autonomous system developed, and was equipped with camera, lidar, and radar sensors. Data collection in Solo was intended to represent the actual environmental conditions on the Bathara Kresna tram line which is mixed traffic, including various vehicle objects, pedestrians, and various signs or slogans used on the tram rail line. Data collection at PT INKA, Madiun, is intended to represent data with actual sensor placement on the tram for which the autonomous system is being developed. These three types of datasets are used to complement online datasets such as Kitti, NuScene, FRSign[9] which represent conditions in developed countries, while the three types of data collected independently represent local data typical of Indonesia. Considering the importance of data obtained from these sensors, the research and innovation team formed a special data engineering team, tasked with managing the data, labeling/annotating it, and preparing the datasets needed to build various types of deep learning models. The availability of good datasets in sufficient quantities is very important in the development of deep learning models[2-4,13,18-20].

Figure 10. Embedded GPU and display system installed in an autonomous tram

The integration of camera sensors, GPS, IMU, and embedded GPU as well as the drive-by-wire system on the electric tram was carried out at PT INKA as shown in Figure 10. Further fine tuning of the algorithm was also carried out at this stage. Specifically, communication was carried out between the embedded GPU and the drive-by-wire system, to ensure that commands given by the Embedded GPU in various types (acceleration, deceleration, buzzer sounding commands and horn sounding commands) could be executed by the tram via the drive-by-wire system. In the first year of this research and innovation, the development of a driver assisted system was carried out. There are several features created, namely 1) object detection using a camera, 2) collision avoidance assist, 3) speed limit assist, 4) driver facial recognition, 5) driver attention warning. While in the second year, 1) object detection using camera-lidar fusion, 2) adaptive cruise control, 3) fully autonomous tram system were developed.

The camera-LiDAR fusion-based object detection method implemented on the Nvidia Drive AGX Pegasus embedded computer combines the CenterPoint [25] and YOLOv3 [26] architectures. Each of these architectures is capable of detecting objects on the road such as car, truck, pedestrian, and cyclist classes. The implementation of the fusion method on the embedded computer uses the C++ programming language using the Velodyne VLP-32C LiDAR hardware and the Sekonix SF3325-100 camera sensor. This fusion program is combined with the main program implemented on the autonomous tram. The output of this fusion method will be used by the decision-making system on the autonomous vehicle to perform actions. Figure 11 below is the Graphical User Interface (GUI) of the implementation on the embedded computer for camera-LiDAR fusion.

Figure 11. Interface display on an embedded computer

As seen in Figure 11, the right side displays a point cloud generated from data from the Velodyne VLP-32C LiDAR sensor. The bottom center displays an image obtained from the Sekonix SF3325-100 camera sensor. The left side displays a map and the vehicle's location, obtained from the GPS sensor.

E. Pengujian

It is important to note that prior to conducting integrated field testing, a simulation environment was developed involving two computing platforms: 1) a PC server with Carla installed, and 2) an Nvidia Embedded GPU. Carla is a simulation and animation environment commonly used in autonomous vehicle development. These two computing platforms were integrated with Carla to display vehicle animations in an urban environment, and an AI pipeline implemented on the Nvidia Embedded GPU. Using this simulation environment, various autonomous tram operating scenarios can be implemented, including situations where real-world testing is not possible due to various limitations.
The performance testing of the camera-LiDAR fusion-based 3D object detection method developed in this study was evaluated on the public KITTI dataset. The KITTI dataset has prepared several sample groups consisting of 7481 training samples and 7518 test samples. Specifically for the test samples, no ground truth labels are available. Therefore, performance calculations are carried out directly by sending detection results to the official KITTI Dataset server. Meanwhile, for the training samples, ground truth labels are available, which are further divided into 3712 training samples and 3769 validation samples.

The camera-LiDAR fusion method developed in this study is capable of fusion to detect 3D objects. By combining several combinations of camera-based 2D object detection methods and 3D object detection methods, the developed fusion method shows improved performance compared to the baseline used when performing fusion. The success of the fusion method tested with several combinations of object detection methods shows that the fusion method architecture can be used modularly or can be combined and separated with various other systems with various types of object detection methods. In general, based on this, the highest detection performance based on camera-LiDAR fusion was obtained when the SECOND-YOLOv3 combination with an mAP of 91.15% for BEV mAP, and 84.74% for 3D mAP.

Testing of the fully autonomous system on a tram was carried out on Jalan Yos Sudarso, Madiun, a significant step in testing and implementing autonomous technology in a real-world environment. The test was conducted by considering various obstacles that may appear on the road, such as pedestrians, motorcycles, and cars, as shown in Figure 12. Meanwhile, the interface of the fully autonomous system is shown in Figure 13. Various obstacles that may be encountered, such as pedestrians suddenly appearing on the tram track, motorcycles maneuvering, and cars interacting with the tram, were all tested to ensure that the autonomous system can respond quickly and effectively. This test aims to validate the autonomous system's ability to detect and avoid these obstacles.

Figure 12. Testing process in a mixed traffic environment
Figure 13. Interface display during fully autonomous system testing

Test results showed that the autonomous system performed well and met expectations. During testing, the tram's braking, acceleration, and speed limit control processes operated efficiently and within predetermined parameters. The system's ability to identify and respond to obstacles, including pedestrians, motorcycles, and cars, provided confidence that the tram's autonomous technology is reliable and safe for implementation in everyday traffic situations. Real-world road testing provided real-world validation of the autonomous system's reliability and safety. The positive results from these tests provide a solid foundation for the continued development and wider implementation of autonomous trams, and make a significant contribution to the development of autonomous technology in the transportation sector.

Figure 14 The results of the fusion method detection were successfully sent to the decision-making system and provided a stop action.

In addition, testing was conducted while the autonomous tram was traveling along the rails. In this test, there would be pedestrians standing still on the rails in front of the autonomous tram. The output of the camera-LiDAR fusion-based object detection perception system was sent to the autonomous tram's decision-making system. The results of this test are shown in Figure 14, which shows that the camera-LiDAR fusion-based object detection perception system successfully detected pedestrians standing still in front of the autonomous tram. The output of this detection result, in the form of the object's position in three-dimensional LiDAR coordinates, was sent to the decision-making system. So that when the object is less than 20 meters away, the autonomous tram will provide emergency braking action or Emergency Braking System (EBS).

F. Benefits, Challenges, and Expectations
There are many benefits to be gained from implementing this technology, including:

  1. Safety and Security: The use of autonomous technology can reduce the risk of accidents caused by human factors. These systems can respond quickly to emergency situations and take necessary actions to prevent accidents.
  2. Transportation Efficiency: Autonomous trams can better manage travel schedules and routes, reducing congestion and waiting times. This will significantly contribute to urban transportation efficiency.
  3. Greenhouse Gas Emission Reduction: By using environmentally friendly technology, autonomous tram systems can help reduce greenhouse gas emissions produced by traditional transportation.
  4. Empowering Local Technology: The collaboration between ITB and PT INKA also encourages the empowerment of local technology in Indonesia. The development and production of autonomous tram systems domestically will create jobs and increase the capacity of the country's technology industry.

While this project holds significant potential, several challenges remain, such as regulatory compliance, public acceptance, and the development of supporting infrastructure. With strong collaboration between academia and industry, and government support, it is hoped that this project can serve as a model for the development of advanced autonomous transportation technology in Indonesia.

G. Kesimpulan
Based on the results of the research that has been carried out, several conclusions can be drawn which are described as follows:

  1. The development of a fully autonomous system architecture with a 3D object detection perception system based on camera and LiDAR fusion using fuzzy algorithms for autonomous vehicles was successfully carried out.
  2. The implementation of the camera-LiDAR fusion method on an embedded computer is able to detect objects around autonomous vehicles well even in dark and rainy lighting conditions, which can be used by decision-making systems on autonomous vehicles.
  3. The collaborative initiative between ITB and PT INKA to develop an autonomous tram system exemplifies how collaboration between academia and industry can create innovative solutions to improve people's quality of life. This system will not only advance the transportation sector but also strengthen Indonesia's position in the global technology arena. Through this joint effort, it is hoped that Indonesia can become a pioneer in the application of autonomous technology in the transportation sector.

H. Ucapan Terima Kasih
We would like to express our sincere gratitude to all parties who have provided significant support and contributions to the smooth development of this autonomous tram system. We extend our gratitude to the Indonesian government, particularly the LPDP (Indonesian Institute of Technology and Development), for its financial support, which has provided a strong foundation for this research. The funding provided by LPDP over three years has been a key driver in realizing this innovation. The trust and support of the government have been key to our success in overcoming the technical and logistical challenges faced in developing this system. We also extend our gratitude to the entire research team, industry partners, and all individuals who have contributed with dedication and enthusiasm. Together, these parties form a solid foundation for advancing future transportation technology.

I. Daftar Pustaka

  1. A. Kim, A., A. Osa, L. Taixee, “EagerMot : 3D Multi Object Tracking via Sensor Fusion”, IEEE International Conference on Robotics and Automation, 2021
  2. Mohammed Elgendy, “Deep Learning for Vision Systems”, Manning, 2020
  3. Francois Chollet, “Deep Learning with Python”, Manning, 2021
  4. Aston Zhang dkk, “Dive into Deep Learning”, 2022 (https://d2l.ai)
  5. C. Di Palma, V. Galdi, V. Calderaro, F. De Luca, F., “Driver Assistance System for Trams: Smart Tram in Smart Cities”, 2020 IEEE International Conference on Environment and Electrical Engineering, 2020
  6. K. Doherty, D. Fourie, J. Leonard, “Multimodal Semantic SLAM with Probabilistic Data Association”, International Conference on Robotics and Automation (ICRA), 2019
  7. Y. Dong, S. Wang, J. Yue, J., C. Chen, S. He, H. Wang, B. He, B., “A Novel Texture-Less Object Oriented Visual SLAM System”, IEEE Transactions on Intelligent Transportation Systems, 2019, 1-14.
  8. D. Frost, V. Prisacariu, D. Murray, “Recovering Stable Scale in Monocular SLAM Using Object-Supplemented Bundle Adjustment”, IEEE Transaction on Robotics, 2018, 34(3), 736-747.
  9. J. Harb, N. RĂ©bĂ©na, R. Chosidow, G. Roblin, R. Potarusov, H. Hajri, “FRSign: A Large-Scale Traffic Light Dataset for Autonomous Trains”, ArXiv preprint arXiv:2002.05665, 2020
  10. H. Liu, “Robot Systems for Rail Transit Applications”, Elsevier, 2020
  11. S. Liu, “Engineering Autonomous Vehicle and Robots: the Dragonfly Modular Based Approach”, Wiley, 2020
  12. I. Rusli, B.R. Trilaksono, W. Adiprawita, “RoomSLAM: Simultaneous Localization and Mapping with Objects and Indoor Layout Structure”, IEEE Access, 2020
  13. S. Raschka, dkk, “Machine Learning with Pytorch and ScikitLearn : Develop Machine Learning and Deep Learning Models with Python”, Packt Publishing, 2022
  14. Siemens, “Siemens Mobility on the Way to an Autonomous Tram”, 2019, https://assets.new.siemens.com/siemens/assets/api/uuid:9d7d12df-a30c-43bda06dbff6dfe96543/presentation-itpautonome-tram-e.pdf
  15. Siemens, “Teaching trams to drive. On the way to smart and autonomous trams: A Siemens Mobility research project, Siemens Mobility GmbH, 2019, https://assets.new.siemens.com/siemens/assets/api/uuid:bc2811c4-3d26-460d-9472-9372d5ce32d7/autonomous-tram.pdf
  16. Systra, “Automated and Autonomous Public Transport : Possibilities, Challenges and Technologies”, 2020, https://www.systra.com/wp-content/uploads/2020/09/systra-automated_and_autonomous_public_transport_2018.pdf
  17. A. Pohan, B.R. Trilaksono, S.P. Santosa, A.S. Rohman, “Path Planning Algorithm Using the Hybridization of Rapidly Random Tree and Ant Colony Systems, IEEE Access, 2021
  18. D. Parekh dkk., “A Review on Autonomous Vehicles : Progress, Methods, and Challenges”, Electronics, 2022, 11(14)
  19. Z. Zhu dkk, “Deep Learning for Autonomous Vehicle and Pedestrian Interaction Safety”, Safety Science, Vol. 145, 2022
  20. J. Ren dkk., “Applying Deep Learning for Autonomous Vehicles : A Survey”, 4th International Conference on Artificial Intelligence and Big Data, 2021
  21. S. Liu dkk, “Creating Autonomous Vehicle Systems”, Morgan & Claypool Publishers, 2020
  22. Alaba, S. Y., & Ball, J. E. (2022). A survey on deep-learning-based lidar 3d object detection for autonomous driving. Sensors, 22(24), 9577
  23. Wang, Z., Wu, Y., & Niu, Q. (2019). Multi-sensor fusion in automated driving: A survey. IEEE Access, 8, 2847-2868.
  24. Pang, S., Morris, D., & Radha, H. (2020). CLOCs: Camera-LiDAR object candidates fusion for 3D object detection. In 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (pp. 10386-10393). IEEE.
  25. Yan, Y., Mao, Y., & Li, B. (2018). Second: Sparsely embedded convolutional detection. Sensors, 18(10), 3337.
  26. Redmon, J., Divvala, S., Girshick, R., & Farhadi, A. (2016). You only look once: Unified, real-time object detection. In Proceedings of the IEEE conference on computer vision and pattern recognition (pp. 779-788).