MTE 实现详解,第一部分:实现测试

访问原始链接 Google 翻译

作者:Mark Brand,Project Zero

背景

2018年,ARM在架构版本v8.5a中提出了一种硬件实现的标记内存方案,称为MTE(内存标记扩展)。

从2022年中到2023年初,Project Zero获得了实现此指令集扩展的预生产硬件,以评估其实现的安全性。我们特别关注的是,是否有可能利用此指令集扩展来实现有效的安全缓解措施,还是其用途仅限于调试/故障检测。

根据v8.5a规范,MTE可以在两种不同的模式下运行,这两种模式基于每个线程进行切换。第一种模式是同步MTE(sync-MTE),在这种模式下,内存访问时的标记检查失败将导致执行该访问的指令在退休时产生故障。第二种模式是异步MTE(async-MTE),在这种模式下,标记检查失败不会直接(在架构层面)导致故障。相反,标记检查失败将设置一个每核心的标志,然后可以从内核上下文轮询该标志,以检测何时发生了无效访问。

这篇博客文章记录了我们迄今为止进行的测试、我们从中得出的结论,以及重复这些测试所需的代码。这些测试旨在探索MTE硬件实现的细节,以及Linux内核中当前对MTE的软件支持状态。所有测试都基于在静态链接的独立二进制文件中手动实现的标记,因此应该很容易在任何兼容硬件上重现这些结果。

术语

在设计实现安全功能时,明确具体的保护目标非常重要。为了在本文其余部分保持清晰,我们将定义一些在讨论此问题时使用的特定术语:

  1. 缓解措施 - 缓解措施是指能够降低某个漏洞或某类漏洞实际可利用性的措施。预期攻击者能够(并且最终会)找到绕过它的方法。例如DEP、ASLR、CFI。

    1. "软"缓解措施 - 如果预期攻击者需要支付一次性成本来绕过缓解措施,我们则认为该缓解措施是"软"的。通常这是针对每个目标的成本,例如开发ROP链来绕过DEP,这通常可以在针对同一目标软件的不同漏洞利用中大量复用。
    2. "硬"缓解措施 - 如果预期攻击者无法开发出可在不同漏洞之间复用的绕过技术(例如,不借助额外的漏洞),我们则认为该缓解措施是"硬"的。例如ASLR,通常要么通过使用单独的信息泄露漏洞来绕过,要么利用同一漏洞开发出信息泄露原语来绕过。
      请注意,上下文可能会影响缓解措施的"硬度"——如果某个代码库中特别富含允许构建信息泄露的代码模式,那么攻击者很有可能开发出一种可靠、可复用的技术,在该代码库内将各种内存破坏原语转化为信息泄露。
  2. 解决方案 - 解决方案是指能够消除某个漏洞或某类漏洞可利用性的措施。预期攻击者绕过解决方案的唯一途径是解决方案中存在非预期的实现缺陷。

例如(基于纯粹理论上的内存标记实现):

  • 为堆分配随机分配标记不能被视为针对任何堆相关内存破坏类别的解决方案,因为这最多只能提供概率性保护。
  • 理论上,为相邻的堆分配分配奇数和偶数标记可能能够为线性堆缓冲区溢出提供解决方案。

硬件实现

我们关于硬件实现的主要问题是:是否存在一种推测性侧信道,可以在无需架构上执行标记检查的情况下,泄露标记检查是否成功?

预计Spectre类型的推测执行侧信道仍然允许攻击者从内存中泄露指针值,从而间接泄露标记。这里我们考虑的是实现是否引入了额外的推测性侧信道,使得攻击者能够更高效地泄露标记检查的成功/失败。

太长不看版:在我们的测试中,我们没有发现额外的1推测性侧信道可以用于执行此类攻击。

1. MTE能阻止Spectre吗?(不能)

MTE能够阻止利用Spectre类型弱点的唯一方式是,让推测执行暂停流水线,直到标记检查完成2。这听起来可能很理想("防止利用Spectre类型弱点"),但事实并非如此——这将创建一个更强的侧信道,使攻击者能够构建标记检查成功的预言机,从而削弱MTE实现的整体安全性。

这很容易测试。如果我们仍然可以在用于推测访问的指针标记不正确的情况下,利用Spectre从推测执行中泄露数据,那么情况就不是这样。

我们为safeside演示程序spectre_v1_pht_sa写了一个小补丁来证明这一点:

mte_device:/data/local/tmp $ ./spectre_v1_pht_sa_mte
Leaking the string: 1t's a s3kr3t!!!
Done!
Checking that we would crash during architectural access:
Segmentation fault

2. 标记检查成功/失败对推测窗口长度有可测量的影响吗?(可能没有)

我们想了解一个更深层次的问题:内存访问后的推测执行长度是否受该访问的标记检查成功或失败的影响?如果是这样,我们或许能够构建一个更复杂的推测性侧信道,并以类似的方式使用它。为了测量这一点,我们需要强制发生错误推测,然后对标记检查可能成功或失败的内存执行读取访问,接着我们需要使用推测性侧信道来测量成功推测执行了多少条后续指令。我们编写并测试了一个独立的工具来测量这一点,可以在speculation_window.c中找到。

该工具的工作原理是在运行时生成一个测试函数的代码,该函数包含可变数量的空操作指令,然后在循环中重复执行此测试函数以训练分支预测器,最后触发错误推测。测试函数如下:

ldr  x0, [x0]         ; 这个加载很慢(*x0 未缓存)
cbnz x0, speculation: ; 在预热期间,这个分支总是被采用
ret
speculation:
ldr  x1, [x1]         ; 这个加载很快(*x1 已缓存)
                      ; 标记检查成功或失败将发生在此次访问期间,
                      ; 但在预热期间,标记检查总是成功的。
orr  x2, x2, x1       ; 这是一个空操作(因为 x1 总是 0),但它
... n times ...       ; 在加载(和空操作)之间保持了数据依赖性,
orr  x2, x2, x1       ; 希望防止过多的重排序。
ldr  x2, [x2]         ; *x2 未缓存,如果之后它被缓存了,
                      ; 那么这条指令(很可能)被执行了。
ret

通过这种方式,我们可以测量在最终加载不再执行之前(即,收到第一次加载的结果,我们意识到错误预测了分支,因此在最终加载执行之前刷新流水线)可以插入多少空操作。

为了减少由于分支预测器行为带来的偏差,测试程序分别针对标记检查成功和标记检查失败的情况运行,并且在这两种情况下运行的执行流是相同的。

为了减少由于CPU电源管理/节流行为带来的偏差,每次运行测试程序时,它收集一组样本然后退出。然后我们重复运行此程序,并通过交错运行标记检查成功和标记检查失败的情况,我们试图减少这种对结果的影响。此外,我们使用标准的Linux CPU缩放接口将核心降频到支持的最低频率,这应该能将CPU节流的影响降到最低。

大多数现代移动设备还具有非对称核心设计,例如4+4、2+2+4或1+3+4核心设计。由于被测设备也不例外,程序需要将自己绑定到特定的核心,并且我们需要为每种核心设计分别运行这些测试。读取虚拟计时器非常慢(相对于内存访问指令),这导致在尝试测量单次缓存命中/单次缓存未命中时会产生噪声数据。在正常的现实环境中,共享内存计时器通常是实现这种粒度级别计时的更好方法,这也是我们过去成功使用过的方法。在这种情况下,由于我们想获取每个核心的数据,很难在所有核心上获得一致性,因此我们采用了放大方法,对多个缓存行进行访问,从而获得了更好的结果。

然而,这种方法增加了测量代码的复杂性,也增加了智能预取或类似功能干扰我们测量的可能性。由于我们试图测试硬件的特性,而不是开发需要及时运行的实用攻击,我们决定将这些风险降到最低,转而收集足够的数据,以便能够透过噪声看到模式。

下面的第一张图显示了每个核心在推测窗口结束附近(因此x轴针对每个核心的锚点不同)的x轴区域(空操作计数)的数据,y轴绘制了我们测量结果导致缓存未命中的观测概率。每张图中绘制了两条线;一条红色线表示"标记检查失败"情况,一条绿色线表示"标记检查通过"情况。我们几乎看不到绿线——两条线重叠得非常紧密,以至于无法区分。

如果我们真的放大"最大"核心的测量结果,可以看到两条线由于噪声而出现偏差:

另一个有趣的点是注意到最大核心的图表顶部概率约为30%;这(可能)不是因为我们命中了缓存,而是因为计时器读取速度太慢,以至于大多数时候我们无法用它可靠地区分缓存未命中和缓存命中。如果我们查看原始延迟测量值的图表,可以看到这一点:

重要的是,失败情况和成功情况的结果之间没有显著差异,这表明没有明显的推测性侧信道可以让攻击者直接构建侧信道标记检查预言机。

软件实现

我们发现了当前围绕MTE使用的软件实现可能导致绕过的几种方式,如果MTE被用作安全缓解措施的话。如下所述,问题1和2是当前实现的特性,可能可以通过内核更改解决,也可能无法解决。它们的范围都有限,并且只适用于非常特定的条件,因此它们可能对MTE作为安全缓解措施的适用性没有特别重大的影响(尽管它们对此类缓解措施的覆盖范围有一些影响)。

问题3是当前Breakpad实现的一个细节,它突显了一类特定的弱点,任何希望使用MTE作为安全缓解措施的应用程序都需要仔细审计其信号处理代码。问题4是一个根本性的弱点,可能无法有效解决。我们不会将此描述为意味着MTE不能用作安全缓解措施,而是将其视为对此类缓解措施有效性声明的限制。

更具体地说,问题3和4都受到一些相当显著的限制,这些限制会根据漏洞发生的具体上下文对内存破坏问题的可利用性产生不同程度的影响。有关此主题的更多讨论,请参见下文。

太长不看版:在我们的测试中,我们发现了软件方面一些需要改进的地方,并证明了在某些情况下可以利用狭窄的"利用窗口"来避免标记检查失败的副作用,这意味着任何基于异步MTE的缓解措施,无论采用何种标记策略,都可能仅限于"软缓解措施"。

1. 系统调用参数的处理 [异步]

Linux内核当前处理异步MTE检查模式存在一个已记录的局限性,即内核不会捕获对用户空间指针的无效访问。这意味着在内核空间访问标记错误的用户空间指针时,系统调用期间不会检测到。

这个限制可能有其技术上的合理原因,但我们计划在未来研究改进这种覆盖范围。

我们提供了一个演示此限制的示例,software_issue_1.c。

mte_enable(false, DEFAULT_TAG_MASK);
uint64_t* ptr = mmap(NULL, 0x1000,
PROT_READ|PROT_WRITE|PROT_MTE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0);
uint64_t* tagged_ptr = mte_tag_and_zero(ptr, 0x1000);
memset(tagged_ptr, 0x23, 0x1000);
int fd = open("/dev/urandom", O_RDONLY);
fprintf(stderr, "%p %p\n", ptr, tagged_ptr);
read(fd, ptr, 0x1000);
assert(*tagged_ptr == 0x2323232323232323ull);
taro:/ $ /data/local/tmp/software_issue_1
0x7722c5d000 0x800007722c5d000
software_issue_1: ./software_issue_1.c:46: int main(int, char **): Assertion *tagged_ptr == 0x2323232323232323ull' failed.

### 2. Handling of system call arguments [SYNC]

The way that the sync MTE check mode is currently implemented in the linux kernel means that kernel-space accesses to incorrectly tagged user-space pointers result in the system call returning EFAULT. While this is probably the cleanest and simplest approach, and is consistent with current kernel behaviour especially when it comes to handling results for partial operations, this has the potential to lead to bypasses/oracles in some circumstances.

We plan to investigate replacing this behaviour with SIGSEGV delivery instead (specifically for MTE tag check failures).

The provided sample, [software_issue_2.c](https://github.com/googleprojectzero/p0tools/blob/master/MTETest/software_issue_2.c) is a very simple demonstration of this situation.

size_t readn(int fd, void* ptr, size_t len) {

char* start_ptr = ptr;

char* read_ptr = ptr;

while (read_ptr < start_ptr + len) {

ssize_t result = read(fd, read_ptr, start_ptr + len - read_ptr);

if (result <= 0) {

return read_ptr - start_ptr;

} else {

read_ptr += result;

}

}

return len;

}

int main(int argc, char** argv) {

int pipefd[2];

mte_enable(true, DEFAULT_TAG_MASK);

uint64_t* ptr = mmap(NULL, 0x1000,

PROT_READ|PROT_WRITE|PROT_MTE, MAP_ANONYMOUS|MAP_PRIVATE, -1, 0);

ptr = mte_tag(ptr, 0x10);

strcpy(ptr, "AAAAAAAAAAAAAAA");

assert(!pipe(pipefd));

write(pipefd[1], "BBBBBBBBBBBBBBB", 0x10);

// In sync MTE mode, kernel MTE tag-check failures cause system calls to fail

// with EFAULT rather than triggering a SIGSEGV. Existing code doesn't

// generally expect to receive EFAULT, and is very unlikely to handle it as a

// critical error.

uint64_t* new_ptr = ptr;

while (!strcmp(new_ptr, "AAAAAAAAAAAAAAA")) {

// Simulate a use-after-free, where new_ptr is repeatedly free'd and ptr

// is accessed after the free via a syscall.

new_ptr = mte_tag(new_ptr, 0x10);

strcpy(new_ptr, "AAAAAAAAAAAAAAA");

size_t bytes_read = readn(pipefd[0], ptr, 0x10);

fprintf(stderr, "read %zu bytes\nnew_ptr string is %s\n", bytes_read, new_ptr);

}

}

taro:/ $ /data/local/tmp/software_issue_2

read 0 bytes

new_ptr string is AAAAAAAAAAAAAAA

read 0 bytes

new_ptr string is AAAAAAAAAAAAAAA

read 0 bytes

new_ptr string is AAAAAAAAAAAAAAA

read 0 bytes

new_ptr string is AAAAAAAAAAAAAAA

read 0 bytes

new_ptr string is AAAAAAAAAAAAAAA

read 0 bytes

new_ptr string is AAAAAAAAAAAAAAA

read 0 bytes

new_ptr string is AAAAAAAAAAAAAAA

read 16 bytes

new_ptr string is BBBBBBBBBBBBBBB

### 3. New dangers in signal handlers [ASYNC].

Since SIGSEGV is a catchable signal, any signal handlers that can handle SIGSEGV become a critical attack surface for async MTE bypasses. [Breakpad](https://www.google.com/url?q=https://www.chromium.org/developers/crash-reports/&sa=D&source=docs&ust=1661952947380022&usg=AOvVaw3HyXTOsL3B1fRSnMg9GdO3)[/Crashpad](https://www.google.com/url?q=https://www.chromium.org/developers/crash-reports/&sa=D&source=docs&ust=1661952947380022&usg=AOvVaw3HyXTOsL3B1fRSnMg9GdO3) at present have a signal handler (ie, installed in all Chrome processes) which allows a trivial same-thread bypass for async MTE as a security protection, which is demonstrated in async_bypass_signal_handler.c / async_bypass_signal_handler.js.

The concept is simple - if we can corrupt any state that would result in the signal handler concluding that a SIGSEGV coming from a tag-check failure is handled/safe, then we can effectively disable MTE for the process.

If we look at the current [breakpad signal handler](https://source.chromium.org/chromium/chromium/src/+/main:third_party/breakpad/breakpad/src/client/linux/handler/exception_handler.cc;l=328) (the crashpad signal handler has [the same design](https://source.chromium.org/chromium/chromium/src/+/main:third_party/crashpad/crashpad/client/crashpad_client_linux.cc;drc=db6f1567b8caa6dacdd0d46b2a7ac60c5b5ddc82;l=201)):

void ExceptionHandler::SignalHandler(int sig, siginfo_t* info, void* uc) {

// Give the first chance handler a chance to recover from this signal

//

// This is primarily used by V8. V8 uses guard regions to guarantee memory

// safety in WebAssembly. This means some signals might be expected if they

// originate from Wasm code while accessing the guard region. We give V8 the

// chance to handle and recover from these signals first.

if (g_first_chance_handler_ != nullptr &&

g_first_chance_handler_(sig, info, uc)) {

return;

}

It's clear that if our exploit could patch g_first_chance_handler_ to point to any function that will return a non-zero value, then this will mean that tag-check failures are no longer caught.

The example we've provided in [async_signal_handler_bypass.c](https://github.com/googleprojectzero/p0tools/blob/master/MTETest/async_signal_handler_bypass.c) and [async_signal_handler_bypass.js](https://github.com/googleprojectzero/p0tools/blob/master/MTETest/async_signal_handler_bypass.js) demonstrates an exploit using this technique against a simulated bug in the [duktape javascript interpreter](https://duktape.org/).

taro:/data/local/tmp $./async_bypass_signal_handler ./async_bypass_signal_handler.js

offsets: 0x26c068 0x36bc80

starting script execution

access

segv handler 11

Async MTE has been bypassed [0.075927734375ms]

done

It seems hard to imagine designing a signal handler that will be robust against all possible attacks here, as most data that is accessed by the signal handler will now need to be treated as untrusted. It's our understanding that Android is considering making changes to the way that the kernel handles these failures to guarantee delivery of these errors to an out-of-process handler, preventing this kind of bypass.

### 4. Generic bypass in multi-threaded environments [ASYNC].

Since async MTE failures are only delivered when the thread which caused the error enters the kernel, there's a slightly more involved (but more generic) bypass available in multi-threaded environments. If we can coerce another thread into performing our next exploit steps for us, then we can bypass the protection this way.

The example we've provided in [async_thread_bypass.c](https://github.com/googleprojectzero/p0tools/blob/master/MTETest/async_thread_bypass.c) and [async_thread_bypass.js](https://github.com/googleprojectzero/p0tools/blob/master/MTETest/async_thread_bypass.js) demonstrates an exploit using this technique against a simulated bug in the duktape javascript interpreter.

taro:/data/local/tmp $ ./async_bypass_thread ./async_bypass_thread.js

thread is running

Async MTE bypassed!

Note that in practice, the technique here would most likely be to have the coerced thread simply install a new signal handler to effectively reimplement 3. above, but the example invokes a shell command as a demonstration.

## How wide are these windows?

So, we've identified two methods by which a compromised thread might avoid the consequences of performing invalid accesses. This was a known and expected limitation of async-MTE-as-a-security-mitigation — but how significant is this?

In the existing (linux) kernel implementation, an outstanding async-MTE failure is delivered as a SIGSEGV when the thread that performed the faulting access transitions to kernel mode. This means that the user-to-kernel transition can be thought of as a synchronisation barrier for tag-check failures. This leads us to our first rule for async-MTE-safe exploitation:

1. We need to complete our exploit (or disable async-MTE) without needing to make a system call from the faulting thread.

This will already pose a significant limitation in some contexts - for example, many network services will have a read/write loop in between commands, so this would mean that the exploit must be completed within the execution of a single command. However, for many other contexts, this is less of an issue - a vulnerability in a complex file format parser, decompression algorithm or scripting engine will likely provide sufficient control between system calls to complete an exploit (as demonstrated above with the duktape javascript engine).

There is then a second requirement that the exploit-writer needs to bear in mind, which is the periodic timer interrupt. In the kernel configuration tested, this was set to CONFIG_HZ_250`,因此我们可以预期每4毫秒发生一次定时器中断,从而得出我们的第二条规则[3](#h.vc7otbihpy9e):

1.  如果我们想要达到可接受的(95%[4](#h.uu3o4hva4l1j))可靠性,我们需要在约0.2毫秒内完成我们的利用。(+0.01毫秒利用运行时间 => -0.25% 可靠性)

[第二部分](/2023/08/mte-as-implemented-part-2-mitigation.html)将继续从更高层面分析各种不同MTE配置在各种用户空间应用上下文中的有效性,以及我们预期对攻击者成本产生何种类型的影响。

#### [1] 已知/预期Spectre类型的攻击可以从内存中泄露指针值。

#### [2] ARM表示情况不应如此(这个潜在的弱点在MTE设计过程的早期就向他们提出过)。

#### [3] 这假设攻击者无法构建侧信道来确定目标进程中何时发生了定时器中断,这在远程媒体解析场景中似乎是合理的假设,但在本地权限提升的上下文中则不那么合理。

#### [4] 95%是一个粗略估计——第二部分更详细地讨论了这里的推理;但请注意,我们认为这是攻击者可能努力达到的粗略参考点,而不是硬性限制。