July 16, 2026
This paper presents SEAGR (Socially and Emotionally Aware Greeting Robot), a robotic greeting framework designed for human–robot interaction environments involving users from diverse cultural backgrounds and different emotional states. Since greeting behaviour strongly influences first impressions, user comfort, and trust, robots operating in public spaces must be able to interact in a socially appropriate and adaptive manner. However, many existing systems still rely on static greeting routines that do not account for cultural variation, emotional context, or interpersonal distance. SEAGR introduces a dual-layer modulation framework in which cultural identity determines the appropriate greeting type, while affective cues influence how that greeting is executed. The system combines context-aware cultural mapping, emotion-based gesture modulation, and proxemic regulation within a unified Sense–Think–Act architecture. A low-cost prototype is implemented using a USB camera, ultrasonic sensor, Arduino-controlled servos, and a laptop-based Python processing system. This work is presented as a system design and proof-of-concept; empirical validation through user studies is explicitly acknowledged as a current limitation and is identified as the primary direction for future work.
Human–Robot Interaction (HRI) is an increasingly important research area as robots enter social environments such as museums, hospitals, airports, and exhibitions. In such settings, robots must interact with humans in socially appropriate and culturally acceptable ways [1]. Among the various forms of interaction, greeting behaviour plays a critical role as it represents the first point of contact between humans and robots, significantly influencing user perception, trust, and willingness to engage further.
Designing socially appropriate greeting behaviour is genuinely difficult. Human greeting practices vary significantly across cultures: a bow is common in Japan, a Namaste gesture is widely used in India, and a handshake is typically expected in many Western contexts [2]. Robots relying on static routines fail to adapt, producing interactions that feel awkward or offensive. Beyond cultural variation, emotional state shapes how a greeting should unfold. Humans adjust their behaviour depending on whether the other person appears relaxed, hurried, or distressed, yet most existing robotic greeting systems do not incorporate affective cues and produce rigid patterns that feel intrusive in real-world settings.
A third important dimension is proxemics — the spatial distance people maintain during social interaction. Research in HRI has consistently shown that initiating interaction from an inappropriate distance undermines user comfort and acceptance even when all other aspects of the greeting are well designed [3]. Despite its importance, proxemic regulation is often treated as a separate engineering problem rather than an integrated component of a broader social greeting framework.
It is worth noting that human customer service workers typically use standardised greetings — “Hello, how can I help you?” — and adapt only after the customer provides context. The key motivation for cultural adaptation in robotics is not that standardised greetings fail entirely, but that robots lack the implicit social awareness, body language reading, and conversational flexibility that allows humans to adapt naturally after the initial contact. A robot that begins with a culturally informed greeting is more likely to establish initial rapport, particularly with users from cultures where greeting rituals carry significant social weight.
To address these interconnected challenges, this paper presents S.E.A.G.R (Socially and Emotionally Aware Greeting Robot), a framework that generates culturally appropriate and emotionally adaptive greeting behaviours in public social environments. The contributions of this work are as follows:
A dual-layer greeting modulation framework separating cultural greeting selection from emotion-based gesture execution, with formal mathematical descriptions of both layers.
A socially aware interaction pipeline integrating cultural mapping, affective motion modulation, and proxemic regulation within a unified Sense–Think–Act architecture.
A proof-of-concept prototype using low-cost sensors and rule-based decision logic, with a structured pathway toward empirical evaluation.
A discussion of ethical considerations including facial detection consent, cultural parody risk, neurodivergent user needs, and the limitations of badge-based identification.
The remainder of this paper is organised as follows. Section 2 reviews related work. Section 3 describes the proposed SEAGR system architecture with formal models. Section 4 presents the hardware and software implementation. Section 5 discusses limitations. Section 6 presents conclusions and future work.
HRI has emerged as an interdisciplinary field combining robotics, AI, psychology, human factors, and social science [4]. Dautenhahn [1] established that robots must incorporate social intelligence to be accepted in everyday environments. Goodrich and Schultz [4] surveyed key interaction dimensions including autonomy, communication modalities, and interaction architecture. Feil-Seifer and Mataric [5] demonstrated that robots provide meaningful value through social interaction, and Kanda et al. [6] showed that combining embodiment, verbal interaction, and social presence enables long-term human–robot partnerships.
Proxemic behaviour is a core requirement for socially acceptable robot interaction. Walters et al. [3] showed that humans maintain different distances from robots than from other humans, requiring HRI-specific proxemic investigation. Takayama and Pantofaru [7] demonstrated that a robot’s motion behaviour strongly shapes user willingness to approach, while Mumm and Mutlu [8] showed that distancing is psychologically mediated by trust. Samarakoon et al. [9] emphasised that proxemic preferences depend on robot appearance, environment, and user characteristics, and Lehmann et al. [10] confirmed that personal space expectations reflect deeper psychological interpretations of the robot as a social entity. Lawrence et al. [11] reinforced that socially acceptable robot behaviour is governed by broader norms around politeness, timing, and situational context.
Cultural differences similarly shape robot interaction. The CARESSES project [12] demonstrated how knowledge-based systems adapt speech and gestures according to user cultural background for elderly care. Papadopoulos et al. [13] reported that culturally competent robots improve engagement in care settings, and Lim et al. [14] argued that culture must be a central HRI design parameter rather than an optional layer. Trovato et al. [2] confirmed that culturally aligned robot greetings improve acceptance and reduce discomfort.
Affective adaptation is equally central. Breazeal [15] established that emotional models make robots more understandable and socially responsive. Stock-Homburg [16] confirmed across two decades of research that emotional expressiveness increases perceived social presence, trust, and acceptance. Kühnlenz et al. [17] showed that emotional adaptation positively influences user behaviour, and Rawal and Stock-Homburg [18] highlighted the practical limitations of real-world emotion recognition, confirming that robust handling of ambiguous facial information remains an open challenge.
Satake et al. [19] showed that appropriate greeting strategies including approach timing and spatial behaviour significantly increase successful engagement in public environments. Urakami and Seaborn [20] argued that nonverbal cues are essential for social competence, and Mutlu et al. [21] demonstrated that gaze cues shape participant roles during robot conversations. A culturally correct greeting may still feel inappropriate if performed too abruptly or at an incorrect distance; robotic greeting must therefore be understood as a multimodal social act involving not only what greeting is selected, but also when, where, and how it is executed.
Recent work by Hussain et al. [22] demonstrated that lightweight computer vision pipelines achieve reliable real-time gesture execution without specialised hardware, a principle directly relevant to SEAGR’s gesture primitives. Elendu et al. [23] highlighted the ethical dimensions of deploying AI in public environments including consent, data privacy, and algorithmic bias considerations that apply directly to systems involving facial scanning and cultural profiling.
Despite these advances, many existing studies treat cultural adaptation, emotion recognition, proxemic behaviour, and nonverbal communication as separate problems rather than integrating them into a unified framework. The present work is motivated by this gap and explicitly acknowledges the limitations of the current proof-of-concept while establishing a clear pathway toward empirical validation.
The proposed SEAGR system follows a Sense–Think–Act architecture that integrates perception, decision-making, and actuation to generate socially compliant greetings in real time. Fig. 1 illustrates the overall decision flow.
The perception module detects user presence, measures interpersonal distance, and extracts basic affective cues. A USB camera identifies the user’s badge and captures coarse facial expression information. An HC-SR04 ultrasonic sensor measures distance to enforce proxemic compliance. The module does not infer cultural identity from visual appearance, deliberately avoiding appearance-based profiling.
The current prototype consists of a laptop-based processing unit with servo-actuated upper-body components capable of head nodding, bowing, and basic arm gestures. It does not resemble a humanoid robot. Robot morphology is known to influence proxemic preferences and user comfort [9], [10], and the distance thresholds are calibrated for this specific form factor. Deployment on platforms of substantially different appearance would require recalibration.
The proxemic gate is defined formally as follows. Let \(d\) denote the measured distance between the robot and the approaching user. The system defines two thresholds: \(d_{\min}\) (minimum comfortable distance) and \(d_{\max}\) (outer boundary of the social interaction zone). The interaction is activated only when:
\[d_{\min} \leq d \leq d_{\max} \label{eq:proxemic95gate}\tag{1}\]
In the current implementation, \(d_{\min} = 0.5\) m and \(d_{\max} = 2.5\) m, consistent with Hall’s [3] social zone definition. If \(d < d_{\min}\), the robot steps back; if \(d > d_{\max}\), no interaction is initiated.
The decision logic layer is a rule-based framework ensuring predictable, explainable, and ethically compliant behaviour. When a user is detected within the social zone, the system attempts badge identification. The badge identifier is matched against a structured database of user-provided metadata including nationality and preferred language.
Formally, let \(\mathcal{C} = \{c_1, c_2, \ldots, c_n\}\) be the set of cultural profiles stored in the attendee database, and let \(b\) denote the scanned badge identifier. The cultural greeting selection function \(\Phi\) is defined as:
\[\Phi(b) = \begin{cases} g_{c_i} & \text{if } b \mapsto c_i \in \mathcal{C} \\ g_{\text{neutral}} & \text{otherwise} \end{cases} \label{eq:greeting95selection}\tag{2}\]
where \(g_{c_i}\) is the culturally aligned greeting motor primitive associated with profile \(c_i\), and \(g_{\text{neutral}}\) is a default neutral greeting (head nod, verbal salutation, conservative distance) applied when no badge match is found. This fallback is consistent with HRI recommendations that robots should default to minimally intrusive behaviour when user context is unavailable [11], [19]. Table 1 shows the nationality-to-greeting mapping, informed by cross-cultural HRI literature [2], [14].
| Nationality | Gesture | Greeting Phrase |
|---|---|---|
| Japan | Bow | “Konnichiwa” |
| India | Namaste | “Namaste” |
| USA | Handshake Invitation | “Hello” |
| Germany | Handshake Invitation | “Guten Tag” |
| France | Nod | “Bonjour” |
| UAE | Verbal Greeting | “As-salamu Alaikum” |
| Brazil | Handshake Invitation | “Olá” |
| Thailand | Wai Gesture | “Sawasdee” |
| Nigeria | Nod | “Hello” |
The emotional modulation framework adjusts how a greeting is performed based on the detected emotional state of the user. Crucially, this module does not change which greeting is selected that is determined solely by Eq. 2 but modifies the execution parameters of the already-selected greeting to improve perceived responsiveness and social sensitivity.
The system estimates coarse emotional states \(e \in \mathcal{E} = \{\text{relaxed}, \text{neutral}, \text{stressed}, \text{unknown}\}\) from facial expression cues using broad categories to maintain robustness under real-world conditions. The final gesture amplitude is computed as:
\[G_{\text{final}} = \alpha(e) \times G_{\text{base}} \label{eq:gesture95modulation}\tag{3}\]
where \(G_{\text{base}}\) is the predefined cultural gesture amplitude and \(\alpha(e) \in (0, 1]\) is an emotion-dependent scaling factor. The scaling factors are defined as:
\[\alpha(e) = \begin{cases} 1.00 & \text{if } e = \text{relaxed} \\ 0.85 & \text{if } e = \text{neutral} \\ 0.65 & \text{if } e = \text{stressed} \\ 0.85 & \text{if } e = \text{unknown (safe default)} \end{cases} \label{eq:alpha}\tag{4}\]
These values were informed by the proxemics and affective HRI literature [7], [17] and represent conservative design choices that prioritise non-intrusiveness. A stressed user triggers reduced amplitude and shorter duration to avoid intrusiveness, while a relaxed user receives the greeting at full amplitude and normal timing. This approach is consistent with findings by Kühnlenz et al. [17] and Stock-Homburg [16] that even coarse affective adaptation meaningfully improves user comfort and perceived social presence.
The gesture movement duration is also modulated according to:
\[T_{\text{final}} = \beta(e) \times T_{\text{base}} \label{eq:duration}\tag{5}\]
where \(T_{\text{base}}\) is the nominal gesture duration and \(\beta(e)\) follows the same mapping as \(\alpha(e)\). The overall greeting response vector is therefore defined as:
\[\mathbf{R} = \left( \Phi(b),\; G_{\text{final}},\; T_{\text{final}},\; v_{\text{tone}} \right) \label{eq:response95vector}\tag{6}\]
where \(v_{\text{tone}}\) is the vocal tone parameter (normal or softened) selected based on the detected emotional state. Algorithm 1 and Table 2 summarise the gesture scaling procedure.
| Detected Emotion | \(\alpha(e)\) | \(\beta(e)\) | Gesture Amplitude |
|---|---|---|---|
| Happy / Relaxed | 1.00 | 1.00 | Full |
| Neutral | 0.85 | 0.85 | Moderate |
| Stressed / Sad | 0.65 | 0.65 | Reduced |
| Unknown | 0.85 | 0.85 | Moderate (safe default) |
Fig. 3 summarises the dual-layer modulation principle. The proxemic gate (Eq. 1 ) activates the pipeline only when the user is within a socially appropriate distance. Layer 1 selects the greeting type via \(\Phi(b)\) (Eq. 2 ), and Layer 2 modulates execution via Eqs. 3 –6 .
The SEAGR prototype uses a modular architecture integrating low-cost sensing devices, an embedded microcontroller, and a laptop-based perception and decision system. The system consists of three layers Perception, Decision Logic, and Actuation operating together as a complete interaction pipeline. The hardware architecture is illustrated in Fig. 4.
The perception subsystem uses a USB camera and an HC-SR04 ultrasonic sensor. The camera performs face detection, badge recognition, and coarse facial expression analysis using the OpenCV Haar Cascade classifier, which identifies faces efficiently on standard hardware without a GPU. A Dlib 68-point landmark detector then extracts facial cues including eye openness, eyebrow position, and mouth shape. A lightweight FER2013-trained classifier estimates the user’s emotional state, mapped to the three broad SEAGR categories defined in Eq. 4 . The ultrasonic sensor provides reliable real-time proxemic distance measurement, compensating for the instability of purely vision-based distance estimation under variable lighting. Table 3 summarises the sensing components.
| Component | Purpose | Key Specification |
|---|---|---|
| USB Camera | Face detection and badge scanning | 720p / 1080p resolution |
| HC-SR04 Ultrasonic Sensor | Distance measurement (\(d\) in Eq. [eq:proxemic95gate]) | Range: 2 cm – 400 cm |
Robot gestures are executed using servo motors connected to the Arduino microcontroller. The mechanical structure supports upper-body movements including head nodding, bowing, and the Namaste gesture, implemented as predefined motion primitives to ensure smooth and repeatable execution. The Arduino converts high-level gesture commands from the laptop into servo angle positions and timing sequences locally, reducing processing load. An audio speaker delivers verbal greeting output in the appropriate language, producing multimodal interaction. This design is supported by Urakami and Seaborn [20] and Kanda et al. [6], who showed that combining gesture, voice, and embodiment significantly improves perceived social competence. Tables 4 and 5 list the actuation components and software stack respectively.
| Component | Function | Interface |
|---|---|---|
| Arduino Uno | Servo motor control | Serial USB |
| Servo Motors | Gesture execution (\(G_{\text{final}}\)) | PWM control |
| Audio Speaker | Voice greeting output | Audio jack / USB |
| Software Component | Function |
|---|---|
| Python | Main system implementation |
| OpenCV | Face detection and image processing |
| PySerial | Laptop-Arduino communication |
| Arduino IDE | Servo motor control firmware |
| JSON / CSV Database | Attendee metadata (\(\mathcal{C}\)) storage |
| Text-to-Speech Engine | Verbal greeting generation |
Attendee metadata is stored in structured JSON or CSV files. When a badge is recognised, the system retrieves the corresponding cultural profile and evaluates \(\Phi(b)\) (Eq. 2
). The laptop communicates with the Arduino via serial interface, sending symbolic gesture commands (G1 Namaste, G2 bow, G3 handshake, G4 head nod). Gesture and speech are synchronised to produce a
complete greeting act, with amplitude and timing determined by \(\mathbf{R}\) (Eq. 6 ). Prior work by Hussain et al. [22] demonstrated that similarly lightweight pipelines achieve reliable real-time gesture execution for robotic arm control, supporting the feasibility of this approach in time-constrained
interaction scenarios.
While the SEAGR framework presents a coherent and interpretable system design, several important limitations must be explicitly acknowledged.
The most significant limitation is the absence of user studies or subjective evaluations. Claims regarding social appropriateness, user comfort, and adaptive effectiveness are grounded in design reasoning and prior literature rather than experimental evidence; the framework cannot yet be considered validated. Future work must conduct user studies across diverse cultural groups, measuring perceived comfort, trust, naturalness, and social appropriateness against a standardised greeting baseline. The badge-scanning mechanism is well-suited to bounded conference environments but does not generalise to museums or public service spaces. A promising extension is a two-step interaction sequence in which the robot begins with a neutral greeting and then adapts based on the user’s verbal or gestural response, preserving the ethical commitment to avoiding appearance-based cultural inference while broadening applicability.
Coarse emotion recognition is inherently uncertain in real-world conditions. The system cannot reliably distinguish a genuinely distressed user from someone whose neutral resting expression appears tense to the FER2013 classifier — a well-documented perceptual limitation in automated emotion recognition systems. To mitigate this, the framework uses only three broad emotional categories and defaults to the neutral profile (\(\alpha = 0.85\)) when confidence is low, prioritising safety over expressiveness. Facial scanning also raises important ethical considerations. Entering a public space should not constitute implied consent to facial detection. Any real-world deployment must include clear prior notification, explicit opt-in consent mechanisms, and the option to interact without facial scanning, aligning with growing ethical frameworks for AI in public spaces [23].
Standard proxemic rules may break down in noisy or crowded environments, and neurodivergent users may adjust interaction distance in ways that conflict with the robot’s thresholds defined in Eq. 1 . Future iterations should incorporate environmental density awareness and allow users to override distance regulation. A non-humanoid robot performing culturally specific gestures such as a bow or Namaste also risks being perceived as patronising rather than respectful. The current system maps nationality to a default greeting type without asking users how they wish to be greeted by a robot specifically. Future versions should allow users to specify greeting preferences during registration, reducing reliance on cultural generalisation. Finally, the scaling factors \(\alpha\) and \(\beta\) (Eq. 4 ) and the proxemic thresholds \(d_{\min}\), \(d_{\max}\) were determined through iterative design reasoning informed by proxemics literature [3], [7] rather than empirical optimisation, and must be validated and refined through controlled user studies before they can be treated as principled design choices.
This paper presented SEAGR, a socially and emotionally aware robotic greeting framework designed as a system design and proof-of-concept for generating culturally appropriate and context-sensitive greeting behaviours in HRI environments. The system introduces a dual-layer modulation approach with formal mathematical descriptions: cultural greeting selection via the function \(\Phi(b)\) (Eq. 2 ), and emotion-based execution modulation via the response vector \(\mathbf{R}\) (Eq. 6 ), integrated with a proxemic gating condition (Eq. 1 ) within a unified Sense–Think–Act architecture. The framework enables robots to produce socially compliant and adaptive greeting behaviours using accessible, low-cost hardware and lightweight software tools.
It is explicitly acknowledged that the current work has not been validated through user studies. The claims made throughout this paper are grounded in design reasoning and the supporting literature rather than empirical measurement. This is the primary limitation of the work and is recognised as the necessary next step before the framework’s effectiveness can be substantiated not merely deferred as future work.
Future work will conduct user studies across diverse cultural groups measuring perceived comfort, social appropriateness, trust, and engagement against standardised greeting baselines. Ethical protocols for facial scanning consent will be developed from the outset. The research will also explore a two-step adaptive interaction model for non-badged environments, empirical calibration of the heuristic parameters \(\alpha\), \(\beta\), \(d_{\min}\), and \(d_{\max}\), and more robust emotion recognition incorporating speech and gaze cues. These developments will progressively move SEAGR from a conceptually grounded proof-of-concept toward a rigorously evaluated and deployable system for multicultural public interaction environments.