Tossaporn (Tree) Saengja

[lec02] 6.566: Computer Systems Security - LXC, gVisor, Firecracker

Lecture 2 ของ MIT 6.566 -- Computer Systems Security พูดถึง OS/VM isolation: ถ้าต้องรันหลาย application บนเครื่องเดียวกัน เราจะจำกัดความเสียหายอย่างไรเมื่อ application ตัวหนึ่งมี bug หรือเป็น malicious code

ตัวอย่างใกล้ตัวคือ application บนโทรศัพท์ เกมไม่ควรสามารถอ่านข้อมูลของ banking app ได้ แม้ว่าทั้งสองจะทำงานอยู่บน hardware และ OS เดียวกัน หลักของ isolation จึงไม่ใช่การทำให้โปรแกรม “ปลอดภัย” แต่เป็นการกำหนดขอบเขตว่า ถ้าโปรแกรมหนึ่งมีปัญหา มันสามารถเข้าถึงหรือสร้างความเสียหายกับส่วนอื่นของระบบได้แค่ไหน

ในระดับพื้นฐาน OS ใช้กลไกอย่าง virtual memory และ page tables เพื่อแยก address space ของแต่ละ process แต่ isolation ต้องครอบคลุม resource อื่นด้วย เช่น files, processes, network และ CPU/memory usage

Lecture เปรียบเทียบแนวทางหลักหลายแบบ ตั้งแต่ Linux process isolation, LXC/container, gVisor ไปจนถึง Firecracker ซึ่งทั้งหมดพยายามหาจุดสมดุลระหว่าง isolation, performance และ compatibility

แนวทางพื้นฐานของ Linux คือ process และ user ID เช่น file permissions กำหนดว่า user ไหนอ่านหรือเขียน file ไหนได้ ปัญหาคือ mechanism นี้เดิมถูกออกแบบมาสำหรับการ share resource ระหว่าง users มากกว่าการสร้างกำแพงแข็ง ๆ ระหว่าง applications

เรื่องยิ่งซับซ้อนเมื่อ process ต้องอ้างถึง resource ของ process อื่น เช่น kill(pid, ...) เพราะ kernel มี shared state จำนวนมาก และ system calls อ้างถึง resource ผ่านชื่อหรือ identifier ต่าง ๆ เช่น PID, file name, IP address และ port การ isolation จึงไม่ใช่แค่แยก memory แต่ต้องควบคุมว่า process “มองเห็น” resource อะไรได้ด้วย

Linux namespaces แก้ปัญหานี้โดยทำให้ process แต่ละกลุ่มเห็น namespace ของ resource ต่างกัน เช่น PID namespace ทำให้ process ใน container เห็นเฉพาะ process บางกลุ่มได้ แนวคิดเดียวกันใช้กับ filesystem และ network ทำให้กลุ่ม process หนึ่งดูเหมือนกำลังทำงานอยู่ใน Linux system ของตัวเอง

ส่วน cgroups ใช้ควบคุมการใช้ resource เช่น CPU, memory, disk I/O และ network I/O โดยตัวมันเองไม่ใช่ security boundary แต่ช่วยป้องกันปัญหาอย่าง process หนึ่งกิน CPU หรือ memory จน process อื่นทำงานไม่ได้

นี่เป็นฐานของ LXC-style container isolation: เตรียม filesystem ของ container, สร้าง namespaces, ตั้ง cgroups, สร้าง virtual network interface แล้วจึงรัน process ข้างใน

ตรงนี้ต้องแยกคำว่า container ออกเป็นสองเรื่อง เรื่องแรกคือ packaging abstraction: มี filesystem image พร้อม libraries, packages และ dependencies แล้วระบุ command ที่ต้องการรัน เช่น application สองตัวสามารถใช้ Python หรือ library คนละ version บนเครื่องเดียวกันได้

อีกเรื่องคือ isolation mechanism ที่ใช้รัน container นั้น ตัว container abstraction เดียวกันไม่จำเป็นต้องใช้ Linux namespaces เสมอไป แต่สามารถรันผ่าน gVisor หรือแม้แต่ VM ได้ด้วย

ข้อดีของ LXC คือ lightweight และ performance ใกล้ native เพราะ application ยังใช้ host Linux kernel โดยตรง แต่จุดนี้ก็เป็นข้อจำกัดด้าน security เช่นกัน เพราะ container ทุกตัว share kernel เดียวกัน

Linux มี system calls มากกว่า 350 ตัว และยังมี interface ซับซ้อนอื่นอย่าง ioctl ถ้ามี vulnerability ใน kernel ที่ code จาก container เข้าถึงได้ attacker อาจ escape ออกจาก isolation boundary ได้

จึงมี mechanism เพิ่มอย่าง seccomp-bpf สำหรับ filter ว่า process เรียก syscall ไหนได้บ้าง แต่ก็มี trade-off: ถ้าปิด syscall มากเกินไป application อาจทำงานไม่ได้ และแม้เปิดเฉพาะ common syscall ก็ไม่ได้แปลว่าจะไม่มี bug อยู่ใน code path เหล่านั้น

อีกแนวทางคือ gVisor ซึ่งสร้างสิ่งที่คล้าย “userspace kernel” ขึ้นมาแทนที่จะปล่อยให้ application ติดต่อ host Linux kernel โดยตรง

component หลักชื่อ Sentry ทำหน้าที่ intercept และ implement Linux syscall interface ใน user space เขียนด้วย Go ดังนั้น syscall จาก application จะไปหา Sentry ก่อน แทนที่จะเข้า host kernel โดยตรง

สำหรับ filesystem gVisor แยก component ชื่อ Gofer ออกมา ซึ่งมีสิทธิ์เข้าถึง host files มากกว่า Sentry วิธีนี้เป็น privilege separation: ต่อให้ Sentry ถูก compromise ก็ไม่ได้หมายความว่าจะได้ arbitrary access ไปยัง host filesystem ทันที

Lecture มี demo ผ่าน runsc ซึ่งเป็น runtime ของ gVisor เมื่อเข้าไปใน container แล้วใช้ uname -a สิ่งที่เห็นจะเป็น kernel environment ที่ gVisor/Sentry นำเสนอ ไม่ใช่ host Linux kernel โดยตรง ขณะเดียวกัน network สามารถถูกแยกออก และ file บางส่วนสามารถถูก explicitly share ผ่าน Gofer ได้

แนวทางนี้ให้ isolation boundary เพิ่มขึ้นและยัง share resource ได้ละเอียดกว่า VM แต่มี cost เพราะทุก syscall ต้องถูก redirect ไปยัง Sentry ทำให้เกิด context switching, data copying และ overhead เพิ่มเติม อีกประเด็นคือ compatibility เพราะ gVisor ต้อง implement Linux syscall behavior ให้ application เชื่อว่ากำลังคุยกับ Linux จริง ๆ

อีกปลายหนึ่งของ spectrum คือ Virtual Machine ซึ่งแทนที่จะจำลอง Linux syscall interface จะให้ guest มี Linux kernel ของตัวเอง

Hardware virtualization และ KVM ช่วยแยก CPU กับ memory ส่วน external I/O ที่ guest เห็นจะเป็น virtual devices เช่น disk, NIC และ serial port ดังนั้น isolation boundary ระหว่าง guest กับ host ไม่ใช่ Linux syscall interface แต่เป็น interface ของ virtual hardware

ข้อดีคือ interface นี้เล็กและ coarse-grained กว่า Linux kernel interface มาก แต่ VM แบบทั่วไปมี overhead ทั้งเรื่อง startup time, memory และ device emulation

Firecracker จึงออกแบบ microVM โดยใช้ KVM สำหรับ virtual CPU และ memory แต่ตัดความซับซ้อนของ VMM แบบ QEMU ออกไปจำนวนมาก รองรับเฉพาะ virtual devices ที่จำเป็น และเขียน VMM ด้วย Rust โดย implementation ที่ lecture อ้างถึงมีขนาดประมาณ 50K lines of code

ความแตกต่างที่สำคัญคือ filesystem interface ด้วย gVisor สามารถให้ guest access file หรือ directory บางส่วนผ่าน Gofer ได้ แต่ Firecracker มอง storage เป็น block device มากกว่าเป็น host files โดยตรง

block device มี interface ค่อนข้างง่าย เช่นอ่านหรือเขียน block ตามหมายเลข ในขณะที่ filesystem มี state และ operation ซับซ้อนกว่า เช่น directories, variable-length files, symlinks, rename และ append การลดความซับซ้อนของ interface ช่วยลด attack surface แต่แลกกับการ share resource ที่หยาบกว่า

Firecracker ยัง sandbox ตัว VMM เองอีกชั้น โดยใช้ chroot, namespaces, separate user ID และ seccomp-bpf ดังนั้นถ้ามี vulnerability ใน VMM ก็ยังมี isolation boundary เพิ่มอีกชั้นหนึ่ง อย่างไรก็ตาม KVM ยังคงเป็นส่วนหนึ่งของ Trusted Computing Base และ bug ใน KVM ก็ยังสามารถทำลาย isolation ของ Firecracker ได้

ภาพรวมจึงไม่ใช่ว่า container กับ VM เป็น abstraction ที่แข่งขันกันตรง ๆ เพราะ container สามารถมองเป็นรูปแบบการ package application—filesystem image บวก command ที่จะ run—แล้วเลือก isolation implementation ข้างใต้ได้หลายแบบ

LXC ใช้ host Linux kernel โดยตรง จึง lightweight และใกล้ native แต่มี kernel attack surface กว้างกว่า

gVisor เพิ่ม userspace kernel อย่าง Sentry ระหว่าง application กับ host kernel ทำให้ isolation แข็งขึ้นและยัง share resource ได้ค่อนข้างละเอียด แต่เพิ่ม syscall overhead

Firecracker ให้ guest Linux kernel จริงบน KVM และ expose virtual hardware ที่เรียบง่ายกว่า ทำให้ isolation boundary แข็งและ performance บางด้านดีกว่า gVisor แต่ resource sharing และ allocation จะ coarse-grained กว่า

ดังนั้นปัญหาหลักของ OS/VM isolation คือการเลือก isolation boundary ที่เหมาะสม พร้อมกับรักษา performance, overhead และ compatibility กับ software ที่มีอยู่ ซึ่งแต่ละ approach เลือก trade-off คนละตำแหน่งบน spectrum นี้

6.566 Lecture 2: OS and VM isolation

#6.566 #summary #thai