Creating Closeness
As a series, Mass Effect has always emphasized the importance of relationships between characters. In some cases, these relationships are friendships. In others, intimacy is involved.
Creating digital characters that feel lifelike and do not present themselves in a way that ruins the player’s suspension of disbelief is a huge challenge.
Creating digital characters capable of communicating strong emotions between each other is even more complex.
The Mass Effect games often show characters touching each other.
Whether this is Shepard pulling an ally out from flaming wreckage or helping them up from a seated position or laying a hand on a friend who has experienced a tragic loss, these moments heighten the realism of the characters and their plight, as well as strengthen the player’s connection to them.
Their emotions feel real.
Beyond that, the close intimate relationships between characters are as much a technical achievement as they are an artistic one.
Characters touching, kissing, or giving each other emotional gazes—these are all actions that we take for granted as being easy, yet within the digital world of a video game, they require state-of-the-art technology and the team’s dedication to complex problem-solving.
These are difficult problems to solve; you need a team that cares about solving them.
Cinematics vs. Real-Time Story Moments
Mass Effect 3 delivers its story through multiple types of sequences. Each type offers its own advantages and costs to the project, and all require a staggering amount of technology and manpower to create:
- Prerendered cinematic scenes (also known as “cutscenes”) are fully animated sequences that are triggered at specific moments during the game. These sequences tend to be highly produced, with custom lighting, music, and visual and auditory effects, and are reserved for major, significant story moments. Prerendered cinematic sequences have the advantage of offering the animators full control over the circumstances during which they will be viewed by the player, which means they can be highly produced and do not need to account for player input that might modify the outcome. These are essentially short, linear movie clips that offer no interactivity.
- Interactive dialogue scenes need to account for a range of variables; for example, which allies did the player bring to the scene, or what choices have they made elsewhere in the game that have an impact on available dialogues? The player can interact with these scenes through the “dialogue-wheel,” which enables them to make significant choices about the content and outcome of a particular conversation. There are hundreds of these scenes in Mass Effect 3
- In-game dialogue scenes are those in which allies make short comments or observations while the player is moving through the world. The player has full control over the scene, and the allies typically offer this ongoing commentary as they follow the player character through the world. These dialogues require the least amount of effort to integrate into the game but are also the least produced.
As a general rule, the more polished and special-case a piece of content is, the more expensive and limited it is in terms of facilitating reusability. The fact that players can often not tell the quality difference between these types of scenes—all of which are seamlessly blended together throughout the game—is a huge testament to the talent of the Mass Effect 3 development team.
Scaling for Variety
For Mass Effect 3, a full-time animator resided within the cinematics design team, their express purpose being to create animations to support the emotional moments.
When a request for a character moment came through from the writing team, the animator would first work with the cinematic designers in reviewing the thousands of existing animations to see if any currently existed that could meet the needs of the scene.
If not, the animator would try blending snippets of existing animations to attempt to reuse content that was readily available.
If the scene needed animation that did not currently exist or could not be assembled out of available pieces, the animator would work to capture the necessary animations during the next motion-capture recording shoot.
This process meant that dialogue-heavy story scenes that were fully interactive—meaning they were not prerendered cinematics, and the player was left fully in control—boasted a healthy amount of animation content captured expressly to service the needs of these scenes.
With over 300 dialogue scenes in Mass Effect 3, there was a real risk that content reuse could lead to a feeling of repetition and “sameness” about the content.
Players are very sensitive to content reuse, and seeing the same animations used over and over again in multiple scenes can make a game feel like a “budget” production.
Using the above process, the team was able to ensure a feeling of purpose-built content having been created for each of the scenes, adding to the quality of the player’s experience.
Lip Synchronization
As with eye movement, ensuring digital characters’ mouths are moving appropriately—known as “lip synchronization,” or, more commonly, “lip-synch”—is a critical part of delivering a believable, high-impact performance.
While the Mass Effect series used specific technology to automatically generate the movement of characters’ lips in real time, there was also an extensive creative process that went into ensuring the technology was tuned to deliver believable, engaging performances.
The original automated technology could bring the lip-synch to an acceptable baseline level, and with the help of the programming team, the bar could be raised to a new level of quality.
In order to improve the lip-synch, the team developed a process inspired by the breakdown of speech in traditional 2-D character animation.
In this type of work, animators pay specific attention to the shape of the mouth as it creates specific vowel and consonant sounds.
When animators look at the sounds coming out of the mouth, the goal is to ensure that the mouth “snaps” to the vowel phoneme sound at the beginning of the word by culling the consonant at the start of the word, with the exception of certain consonants like M, B, and P.
These vowel phonemes at the beginning of words tend to be where the shape of speakers’ mouths are most exaggerated and is what we tend to read when we are looking at someone’s lips as they are speaking to us.
The second improvement came about because, although the automatic lip-synch technology did a good job, it was not able to distinguish between which phonemes the viewer’s brain would ignore when reading the lips; the tech simply generated individual mouth shapes for all phonemes it detected and for each word in a line of dialogue.
This often resulted in very unnatural, jarring mouth movements.
Another evolution of the tech used in Mass Effect 3 meant that the technology was able to detect instances of “jabber jaw,” those cases where the automatic lip-synchronization tech was generating phoneme shapes at too high a frequency, leaving characters’ lips looking like they were mumbling or teeth like they were chattering.
The final step the animators took to ensure the lip-synch was to the highest quality was another process derived from traditional 2-D animation.
Typically, in games, the lips are animated to be precisely synchronized with the dialogue audio they accompany. In 2-D animation, the lips are animated so that the animation triggers slightly before the viewer hears the audio.
This accounts for how the eye and ear each receive data and how the brain processes this data.
The Mass Effect 3 animation team was able to recreate this subtle but noticeable shift, producing lip-synch results that are highly realistic and can satisfy the ultraperceptive human eye.
Through trial-and-error and a lot of iteration, the animation team was able to use a combination of technical horsepower, creative skill, and raw attention to detail to produce nearly perfect lip-synchronization for the dialogue sequences throughout the game.