Formatter Bytecode#

Background#

LLDB provides rich customization options to display data types (see Variable Formatting). To use custom data formatters, developers need to edit the global ~/.lldbinit file to make sure they are found and loaded. In addition to this rather manual workflow, developers or library authors can ship ship data formatters with their code in a format that allows LLDB automatically find them and run them securely.

An end-to-end example of such a workflow is the Swift DebugDescription macro (see https://www.swift.org/blog/announcing-swift-6/#debugging ) that translates Swift string interpolation into LLDB summary strings, and puts them into a .lldbsummaries section, where LLDB can find them.

This document describes a minimal bytecode tailored to running LLDB formatters. It defines a human-readable assembler representation for the language, an efficient binary encoding, a virtual machine for evaluating it, and format for embedding formatters into binary containers.

Goals#

Provide an efficient and secure encoding for data formatters that can be used as a compilation target from user-friendly representations (such as DIL, Swift DebugDescription, or NatVis).

Non-goals#

While humans could write the assembler syntax, making it user-friendly is not a goal. It is meant to be used as a compilation target for higher-level, language-specific affordances.

Design of the virtual machine#

The LLDB formatter virtual machine uses a stack-based bytecode, comparable with DWARF expressions, but with higher-level data types and functions.

The virtual machine has two stacks, a data and a control stack. The control stack is kept separate to make it easier to reason about the security aspects of the virtual machine.

Data types#

All objects on the data stack must have one of the following data types. These data types are “host” data types, in LLDB parlance.

  • String (UTF-8)

  • Integer (arbitrary precision, always signed)

  • Int (64 bit) (deprecated: use Integer)

  • UInt (64 bit) (deprecated: use Integer)

  • Object (Basically an SBValue)

  • Type (Basically an SBType)

  • Selector (One of the predefine functions)

  • Dictionary (A mutable mapping from String keys to values of any data type)

Object and Type are opaque, they can only be used as a parameters of call.

Instruction set#

Stack operations#

These instructions manipulate the data stack directly.

Opcode

Mnemonic

Stack effect

0x00

dup

(x -> x x)

0x01

drop

(x y -> x)

0x02

pick

(x ... UInt -> x ... x)

0x03

over

(x y -> x y x)

0x04

swap

(x y -> y x)

0x05

rot

(x y z -> z x y)

Control flow#

These manipulate the control stack and program counter. Both if and ifelse expect an Integer or UInt at the top of the data stack to represent the condition.

Opcode

Mnemonic

Description

0x10

{

push a code block address onto the control stack

–

}

(technically not an opcode) syntax for end of code block

0x11

if

(Integer|UInt -> ) pop a block from the control stack, if the top of the data stack is nonzero, execute it

0x12

ifelse

(Integer|UInt -> ) pop two blocks from the control stack, if the top of the data stack is nonzero, execute the first, otherwise the second.

0x13

return

pop the entire control stack and return

Literals for basic types#

Opcode

Mnemonic

Description

0x20

123u

( -> UInt) push an unsigned 64-bit host integer (deprecated: use lit_integer)

0x21

123

( -> Int) push a signed 64-bit host integer (deprecated: use lit_integer)

0x22

"abc"

( -> String) push a UTF-8 host string

0x23

@strlen

( -> Selector) push one of the predefined function selectors. See call.

0x24

123

( -> Integer) push an arbitrary precision signed integer

Conversion operations#

Opcode

Mnemonic

Description

0x2a

as_int

( UInt -> Int) reinterpret a UInt as an Int (deprecated)

0x2b

as_uint

( Int -> UInt) reinterpret an Int as a UInt (deprecated)

0x2c

is_null

( Object -> UInt ) check an object for null (object ? 0 : 1)

Arithmetic, logic, and comparison operations#

Every Integer value on the data stack is signed. +, -, *, /, %, =, !=, <, >, =<, >= are defined for Integer (and deprecated Int/UInt) and operate on the operands’ mathematical values.

<<, >>, &, |, ^, ~ are bitwise operations and operate on an Integer’s underlying two’s complement bit pattern rather than its mathematical value. Because a bitwise operation never treats its operands as having a sign, >> is always a logical (zero-filling) shift, not an arithmetic shift, and a negative operand is not an error.

Opcode

Mnemonic

Stack effect

0x30

+

(x y -> [x+y])

0x31

-

etc …

0x32

*

0x33

/

0x34

%

0x35

<<

0x36

>>

0x40

~

0x41

|

0x42

^

0x50

=

0x51

!=

0x52

<

0x53

>

0x54

=<

0x55

>=

Function calls#

For security reasons the list of functions callable with call is predefined. The supported functions are either existing methods on SBValue, or string formatting operations.

Opcode

Mnemonic

Stack effect

0x60

call

(Object argN ... arg0 Selector -> retval)

Method is one of a predefined set of Selectors.

Sel.

Mnemonic

Stack Effect

Description

0x00

summary

(Object @summary -> String)

SBValue::GetSummary

0x01

type_summary

(Object @type_summary -> String)

SBValue::GetTypeSummary

0x10

get_num_children

(Object @get_num_children -> Integer)

SBValue::GetNumChildren

0x11

get_child_at_index

(Object Integer @get_child_at_index -> Object)

SBValue::GetChildAtIndex

0x12

get_child_with_name

(Object String @get_child_with_name -> Object)

SBValue::GetChildMemberWithName

0x13

get_child_index

(Object String @get_child_index -> Integer)

SBValue::GetChildIndex

0x14

get_parent

(Object @get_parent -> Object)

SBValue::GetParent

0x15

get_type

(Object @get_type -> Type)

SBValue::GetType

0x16

get_template_argument_type

(Type Integer @get_template_argument_type -> Type)

SBValue::GetTemplateArgumentType

0x17

cast

(Object Type @cast -> Object)

SBValue::Cast

0x18

get_synthetic_value

(Object @get_synthetic_value -> Object)

SBValue::GetSyntheticValue

0x19

get_non_synthetic_value

(Object @get_non_synthetic_value -> Object)

SBValue::GetNonSyntheticValue

0x20

get_value

(Object @get_value -> Object)

SBValue::GetValue

0x21

get_value_as_unsigned

(Object @get_value_as_unsigned -> Integer)

SBValue::GetValueAsUnsigned

0x22

get_value_as_signed

(Object @get_value_as_signed -> Integer)

SBValue::GetValueAsSigned

0x23

get_value_as_address

(Object @get_value_as_address -> Integer)

SBValue::GetValueAsAddress

0x24

clone

(Object String @clone -> Object)

SBValue::Clone

0x25

get_pointee_type

(Type @get_pointee_type -> Type)

SBType::GetPointeeType

0x26

get_byte_size

(Type @get_byte_size -> Integer)

SBType::GetByteSize

0x27

create_child_at_offset

(Object String Integer Type @create_child_at_offset -> Object)

SBValue::CreateChildAtOffset

0x40

read_memory_byte

(UInt @read_memory_byte -> UInt)

Target::ReadMemory

0x41

read_memory_uint32

(UInt @read_memory_uint32 -> UInt)

Target::ReadMemory

0x42

read_memory_int32

(UInt @read_memory_int32 -> Int)

Target::ReadMemory

0x43

read_memory_uint64

(UInt @read_memory_uint64 -> UInt)

Target::ReadMemory

0x44

read_memory_int64

(UInt @read_memory_int64 -> Int)

Target::ReadMemory

0x45

read_memory_address

(UInt @read_memory_uint64 -> UInt)

Target::ReadMemory

0x46

read_memory

(UInt Type @read_memory -> Object)

Target::ReadMemory

0x50

fmt

(String arg0 ... @fmt -> String)

llvm::format

0x51

sprintf

(String arg0 ... sprintf -> String)

sprintf

0x52

strlen

(String strlen -> Integer)

strlen in bytes

Dictionary objects#

Dictionary objects are key-value containers, with String value keys, and values of any data type. Dictionary is a reference type, mutating it through one reference is visible through any other reference to the same dictionary (e.g. one obtained earlier with dup). Empty Dictionary objects are created with dict. Dictionary objects are populated with dict_set. Values are retrieved with dict_get. When dict_get is called with a key that is not present in the dictionary, an error is emitted. Use dict_has first to check for a key’s existence. Dictionary operations consumes the Dictionary argument, so dup it first if the Dictionary is needed afterward. For example, to set multiple keys in a row:

dict dup "a" 1 dict_set dup "b" 2 dict_set

A Dictionary may be stored as a value in another Dictionary, but dict_set emits an error if doing so would make a dictionary contain itself, directly or through nested dictionaries.

Opcode

Mnemonic

Stack effect

0x70

dict

( -> Dictionary) create an empty dictionary

0x71

dict_set

(Dictionary String x -> ) set a key to a value

0x72

dict_get

(Dictionary String -> x) look up the value for a key

0x73

dict_has

(Dictionary String -> Integer) check whether a key is present

Byte Code#

Most instructions are just a single byte opcode. The only exceptions are the literals:

  • String: Length in bytes encoded as ULEB128, followed length bytes

  • Int: LEB128

  • UInt: ULEB128

  • Integer: LEB128, sign-extended to a signed value of at least 64 bits

  • Selector: ULEB128

Embedding#

Expression programs are embedded into an .lldbformatters section (an evolution of the Swift .lldbsummaries section) that is a dictionary of type names/regexes and descriptions. It consists of a list of records. Each record starts with the following header:

  • Version number (ULEB128)

  • Remaining size of the record (minus the header) (ULEB128)

The version number is increased whenever an incompatible change is made, either to the layout of the record, or to the formatter ABI (see Calling conventions). Adding new opcodes or selectors is not an incompatible change since consumers can unambiguously detect this and report an error.

Space between two records may be padded with NULL bytes.

A record consists of a dictionary key, which is a type name or regex.

  • Length of the key in bytes (ULEB128)

  • The key (UTF-8)

A regex has to start with ^, which is part of the regular expression.

After this comes a flag bitfield, which is a ULEB-encoded lldb::TypeOptions bitfield.

  • Flags (ULEB128)

This is followed by one or more dictionary values that immediately follow each other and entirely fill out the record size from the header. Each expression program has the following layout:

  • Function signature (1 byte)

  • Length of the program (ULEB128)

  • The program bytecode

Calling conventions#

A record’s version number (see Embedding) also determines the calling convention of its methods. This includes how self (aka this) is represented. This section describes version 2; see Version 1 for the differences in version 1.

The runtime owns self, a Dictionary that starts out empty. The runtime passes a reference to self as the first argument to every method except @summary, followed by that method’s arguments (if any). Methods may modify self in place, and do not return it. The runtime’s ownership of self ensures modifications to the Dictionary are visible to subsequent method calls.

The possible function signatures are:

Signature

Mnemonic

Stack Effect

0x00

@summary

(Object -> ... String)

0x01

@init

(Dictionary Object -> ...)

0x02

@get_num_children

(Dictionary -> ... Integer)

0x03

@get_child_index

(Dictionary String -> ... Integer)

0x04

@get_child_at_index

(Dictionary Integer -> ... Object)

0x05

@get_value

(Dictionary -> ... String)

0x06

@update

(Dictionary -> ... Integer)

The @init method must only be used for one time setup work. Any computation that needs to be reperformed should happen in @update. If not specified, initialization will save the given Object to the self dictionary using the idiomatic key "valobj".

The @update method performs computation that may change over the course of a value’s lifetime, such as interpreting a value’s state, and (re)computing children. The return value is an Integer, where 1 means the previously computed children can be reused, and 0 means they must be refetched.

While it is more efficient to store multiple programs per type key, this is not a requirement. LLDB will merge all entries. If there are conflicts the result is undefined.

Execution model#

Execution begins at the first byte in the program. The program counter of the virtual machine starts at offset 0 of the bytecode and may never move outside the range of the program as defined in the header. The data stack starts with the method’s arguments, as listed in the signature table in Calling conventions.

Error handling#

Errors are unrecoverable, the entire expression will fail if any kind of error is encountered.

Version 1#

Version 1 records use the same record layout as version 2 (see Embedding), but an older calling convention:

  • There is no self dictionary. Instead, self is whatever is left on the data stack after @init and @update run. Nothing constrains its shape: it can be zero, one, or many values (written Object+ below), and it is passed to the other methods in place of the Dictionary.

  • If @init is not specified, self is the given Object.

  • The @update reply is optional. If @update does not leave an Integer 0 or 1 (or a UInt or Int 0 or 1) on top of the data stack, the children are refetched.

  • Selectors use the deprecated UInt and Int types in place of Integer: get_num_children, get_child_index, get_value_as_unsigned, get_value_as_address, and strlen return a UInt, get_value_as_signed returns an Int, and get_child_at_index and get_template_argument_type take a UInt index.

The version 1 function signatures are:

Signature

Mnemonic

Stack Effect

0x00

@summary

(Object -> String)

0x01

@init

(Object -> Object+)

0x02

@get_num_children

(Object+ -> UInt)

0x03

@get_child_index

(Object+ String -> UInt)

0x04

@get_child_at_index

(Object+ UInt -> Object)

0x05

@get_value

(Object+ -> String)

0x06

@update

(Object+ -> Object+ [Integer])