Challenges for machine interpreting in high-risk scenarios

· AI

Challenges for machine interpreting in high-risk scenarios

Discussions on the impact of translation and interpreting technologies in the field of multilingual communication are more important today than ever. Technology is improving rapidly and will continue to do so in the years to come. Consequently, its use in a multitude of contexts is bound to increase. While in most cases the use of speech technologies will become a simple and worries free commodity, there are areas, such in highly regulated markets, where their use might be classified as high-risky.

In the context of live speech translation, high-risk scenarios refer to situations where the accuracy and reliability of communication are critical, and any misinterpretation or misunderstanding could have significant and potentially severe consequences. These scenarios typically involve areas such as:

In these scenarios, the stakes are high because errors in translation can lead to legal injustices, medical errors, loss of life, diplomatic tensions, or other serious outcomes. It must be noted that stakes are equally high when there is no translation available.

The above considerations apply to any kind of translation agent, whether a human or a computer. While there are some – albeit very general – provisions regarding human interpreters, the novelty of speech technologies has not allowed stakeholders enough time to reflect and define the appropriate use of such technology. A round of discussions, research, and applications is necessary. What do decision-makers need to know and what steps must be taken to ensure that speech technologies, such as machine interpreting, are used effectively and responsibly?

Here are some key points that stakeholders will need to keep in mind when starting to discuss this topic, followed by three simple principles that may guide the work ahead of us in the years to come.

Technological Evolution: The shift from rule-based to neural machine translation has significantly improved the accuracy and capacity of machine translation systems. Similarly, speech recognition technologies have evolved to better transcribe speech into text, in an increasing variety of languages and dialects. The advent of Large Language Models is again pushing technology further, and quality of speech translation systems is about to increase in the years to come, for example thanks to emerging abilities such as grounding translation in the communicative event, understanding cultural references, etc.

Fig 1: Speech Translation System

Practical Applications: Speech translation systems are used by the general public to overcome language barriers while traveling, listening to podcasts, etc. In sensitive and professional settings, speech translation systems are increasingly used due to a shortage of human translators and interpreters, especially for, but not limited to, less common languages, or to reduce the costs and increase availability of translation services. These technologies are also utilized in hospitals and police stations, to name just a few, to fill these gaps and to provide basic language support. These might be obviously considered high-risk scenarios, depending on the scope of the conversation. There is a difference, for example, if a person is requesting an appointment or if a doctor-patient consultation needs to be live translates. The implications for the use of such technologies in potentially critical cases are many, and mostly unexplored.

Legal and Ethical Concerns: The use of speech technologies in high-risk scenarios raises several legal and and ethical challenges:

Challenges in regulated markets: Regulated markets require systems and solutions to meet minimum standards through certifications, examinations, or other means. This applies both to human labor as well as to machine translation systems. Unlike many devices used in the judiciary or hospitals that undergo rigorous testing and vetting, there is still no certification available for this emerging technology. At the moment of speaking we lack sufficient knowledge about how and if it is possible to provide clear error margins or accuracy thresholds for speech translation systems.

Principles to focus on in the coming years

While the use of speech technology in non-regulated markets will be shaped by its adopters and users based on their needs, perceptions, and evaluation criteria, responsible adoption in high-risk scenarios and regulated markets will require profound knowledge, mutual agreements, and regulations. Many discussions need to be initiated, and significant work lies ahead. Here, I identify three high-level principles, or calls to action, for the discussions to come:

While speech translation aims to offer a scalable accessibility solution, particularly in non-critical situations, significant challenges arise in high-stakes scenarios and regulated markets. To mitigate risks and enhance the benefits of automatic speech translation in such critical use cases, substantial efforts are required. A community of experts must be trained, and new knowledge must be generated to address these challenges effectively.

Image by https://st-benchmark.github.io