Amazon SDE 面经(2024 New Grad)
岗位:Software Development Engineer, New Grad(Seattle)
背景:美硕 CS Top 20,一段 AWS 暑期实习(return offer 面试)
结果: offer,评级 L4(New Grad 标准)
时间线:7.15 投递 → 7.20 OA → 8.5 VO Schedule → 8.20 Round 1 → 8.22 Round 2 → 9.1 offer
面试流程概览
Amazon 的 New Grad 面试流程相对标准化:
| 阶段 | 内容 | 时长 |
|---|---|---|
| OA(Online Assessment) | 2 道算法 + 工作风格问卷 | 105min |
| VO Round 1 | 2 道 LP + 1 道 Coding | 60min |
| VO Round 2 | 2 道 LP + 1 道 Coding | 60min |
| Debrief | 两位面试官共同讨论 | 内部流程 |
没有系统设计轮次,这是 New Grad 与 Experienced Hire 的主要区别。
OA(Online Assessment)
算法部分(2 题,每题 70min)
Amazon 的 OA 由 Hackerrank 提供,两道算法题难度通常是 Easy + Medium。
我的题目回忆:
题目 1:Reorder Data in Log Files(LeetCode 937,Easy)
def reorderLogFiles(logs):
def sort_key(log):
identifier, rest = log.split(" ", 1)
return (0, rest, identifier) if rest[0].isalpha() else (1,)
return sorted(logs, key=sort_key)
题目 2:Number of Islands(LeetCode 200,Medium)
def numIslands(grid):
if not grid:
return 0
rows, cols = len(grid), len(grid[0])
count = 0
def dfs(r, c):
if r < 0 or c < 0 or r >= rows or c >= cols or grid[r][c] != '1':
return
grid[r][c] = '0'
dfs(r+1, c)
dfs(r-1, c)
dfs(r, c+1)
dfs(r, c-1)
for r in range(rows):
for c in range(cols):
if grid[r][c] == '1':
count += 1
dfs(r, c)
return count
OA 建议:
- 两道题建议在 40-60 分钟内完成,留出时间检查边界条件
- 编译器只有少量测试用例,需要自己多想 edge case
- 不需要提交即最优解,但要通过所有 visible test cases
工作风格问卷(Work Style Assessment)
约 30 分钟,没有选择对错,全是情景判断。例如:
“When you encounter a problem you’ve never seen before, you usually:”
A. Research it on your own for a while before asking
B. Ask a teammate immediately to save time
C. Try a quick fix and see if it works
D. Document the problem first before doing anything
答题策略:
- 仔细阅读 Amazon 的 16 条 LP,理解其价值观
- Customer Obsession 和 Ownership 是最重要的两条
- 不要选极端答案,选择体现"先尝试再求助"、“团队协作”、“以客户为中心"的选项
VO Round 1(LP + Coding)
面试官是一位 L5 SDE,在 AWS S3 团队工作了 6 年。开场白很标准:“I’ll ask a few behavioral questions first, then we’ll do coding.”
LP 问题 1:Customer Obsession(10min)
问题:Tell me about a time you went above and beyond for a customer.
我的回答(STAR):
S:实习期间负责一个数据分析工具,用户是内部的金融分析师团队。
T:一位分析师反馈说工具导出的 CSV 在 Excel 中打开时中文字符乱码。
A:
- 我不仅修复了编码问题(从 UTF-8 改为 UTF-8 BOM),还主动调研了 Excel 的编码兼容性;
- 添加了导出格式选项(CSV/Excel/XLSX),让用户自己选择;
- 编写了 FAQ 文档,包含常见编码问题的解决方案;
- 在团队周会上分享了这个 issue,推动我们在所有导出功能中统一处理编码。
R:该分析师后续在团队满意度调查中给了满分,她的团队有 3 个人开始高频使用我们的工具。
追问:
- “What would you do if fixing this encoding issue took 3 days instead of 3 hours?”
- “How did you balance this unplanned work with your sprint commitments?”
LP 问题 2:Ownership(10min)
问题:Tell me about a time you took on something outside your responsibility.
我的回答:
S:实习团队的 on-call runbook 非常陈旧,很多排查步骤已经失效。
T:维护 runbook 本来不是实习生的职责,但 on-call 时经常因为文档错误浪费大量时间。
A:
- 我在两次 on-call 中记录了所有文档错误和缺失的步骤;
- 用周末时间重写了 runbook,添加了故障排查决策树;
- 邀请团队成员 review,根据反馈补充了常见错误案例;
- 将 runbook 从内部 wiki 迁移到 searchable 的 documentation 平台。
R:后续 on-call 的平均 MTTR(Mean Time To Recovery)从 45 分钟降到 15 分钟。
算法题(35min)
题目:Course Schedule(LeetCode 207,Medium)
判断有向图中是否有环(拓扑排序)。
from collections import deque, defaultdict
def canFinish(numCourses, prerequisites):
# 构建图和入度数组
graph = defaultdict(list)
in_degree = [0] * numCourses
for course, prereq in prerequisites:
graph[prereq].append(course)
in_degree[course] += 1
# Kahn's algorithm
queue = deque([i for i in range(numCourses) if in_degree[i] == 0])
processed = 0
while queue:
course = queue.popleft()
processed += 1
for next_course in graph[course]:
in_degree[next_course] -= 1
if in_degree[next_course] == 0:
queue.append(next_course)
return processed == numCourses
追问:
- 时间/空间复杂度? $O(V + E)$ 时间,$O(V + E)$ 空间
- 如果要求返回一种有效的课程顺序? 用同样的 Kahn 算法,队列中弹出的顺序就是拓扑序
- 如果用 DFS 怎么做? 维护三种状态(未访问/访问中/已访问),DFS 遇到「访问中」的节点说明有环
def canFinishDFS(numCourses, prerequisites):
graph = defaultdict(list)
for course, prereq in prerequisites:
graph[prereq].append(course)
# 0=未访问, 1=访问中, 2=已访问
state = [0] * numCourses
def dfs(course):
if state[course] == 1: # 访问中,有环
return False
if state[course] == 2: # 已访问
return True
state[course] = 1
for next_course in graph[course]:
if not dfs(next_course):
return False
state[course] = 2
return True
for i in range(numCourses):
if not dfs(i):
return False
return True
VO Round 2(LP + Coding)
面试官是一位 L6 Senior SDE,在 Alexa 团队。这一轮的 LP 追问明显更深。
LP 问题 1:Dive Deep(10min)
问题:Tell me about a time you uncovered a problem by digging into the data.
我的回答:
S:实习期间负责一个批处理任务,每天凌晨处理前一天的日志数据。
T:某天早上收到告警说任务比平时多跑了 2 小时。
A:
- 我先看 CPU 和内存监控,发现内存使用飙升到 90%;
- 深入排查发现是某个新加的聚合逻辑在极端 case 下内存泄漏;
- 进一步分析日志发现,泄漏只发生在 “user_id = null” 的脏数据上;
- 修复了空值处理逻辑,并添加了数据质量检查;
- 在 pipeline 入口处增加了数据校验,防止脏数据进入。
R:任务运行时间恢复稳定在 30 分钟内,后续 3 个月没有再出现类似告警。
追问:
- “How did you know it was a memory leak and not just high memory usage?”
- “What metrics did you look at first, and why that order?”
- “If the issue only happened in production and you couldn’t reproduce it locally, what would you do?”
LP 问题 2:Have Backbone; Disagree and Commit(10min)
问题:Tell me about a time you disagreed with your team and how you handled it.
我的回答:
S:团队讨论是否在一个高并发接口中引入 Redis 缓存。
T:我的导师( senior engineer )认为不需要缓存,因为当前 QPS 不高;但我认为这个接口的增长趋势很快,应该提前优化。
A:
- 我没有直接反对,而是花了一个下午收集了历史 QPS 增长数据(过去 6 个月增长了 300%);
- 做了一个简单的压测,模拟 10 倍 QPS 下的响应时间;
- 在团队会议上展示了数据,提出「先实现但默认关闭,通过 feature flag 控制」的中立方案;
- 最终团队同意添加缓存层,但只在监控指标超过阈值时自动开启。
R:一个月后 QPS 确实翻倍,缓存自动启用,接口 P99 保持在 50ms 以下。我的导师在周会上特别肯定了我的 proactive 态度。
算法题(35min)
题目: LRU Cache(LeetCode 146,Medium)
这道题在 Amazon 面试中极其高频。
class DLinkedNode:
def __init__(self, key=0, value=0):
self.key = key
self.value = value
self.prev = None
self.next = None
class LRUCache:
def __init__(self, capacity: int):
self.capacity = capacity
self.cache = {}
self.head = DLinkedNode()
self.tail = DLinkedNode()
self.head.next = self.tail
self.tail.prev = self.head
self.size = 0
def get(self, key: int) -> int:
if key not in self.cache:
return -1
node = self.cache[key]
self._move_to_head(node)
return node.value
def put(self, key: int, value: int) -> None:
if key in self.cache:
node = self.cache[key]
node.value = value
self._move_to_head(node)
else:
node = DLinkedNode(key, value)
self.cache[key] = node
self._add_to_head(node)
self.size += 1
if self.size > self.capacity:
removed = self._remove_tail()
del self.cache[removed.key]
self.size -= 1
def _add_to_head(self, node):
node.prev = self.head
node.next = self.head.next
self.head.next.prev = node
self.head.next = node
def _remove_node(self, node):
node.prev.next = node.next
node.next.prev = node.prev
def _move_to_head(self, node):
self._remove_node(node)
self._add_to_head(node)
def _remove_tail(self):
node = self.tail.prev
self._remove_node(node)
return node
追问:
- 如何实现线程安全? 加锁(
threading.Lock())或者用collections.OrderedDict(Python 3.7+ dict 有序) - 如果缓存容量是 1GB,但实际数据只有 100MB,如何优化内存使用? 可以使用更紧凑的数据结构,或者将不常用数据序列化到磁盘
- LRU 的淘汰策略有什么缺点? 对突发热点不友好(scan 攻击),可以改进为 LRU-K 或 LFU
面试复盘与建议
Amazon 面试的核心竞争力
| 维度 | 权重 | 建议 |
|---|---|---|
| LP | ⭐⭐⭐⭐⭐ 50% | 准备 15-20 个 STAR 故事,每条 LP 至少覆盖 2 次 |
| 算法 | ⭐⭐⭐⭐ 35% | LeetCode Easy-Medium 为主,代码一次过 |
| 沟通 | ⭐⭐⭐ 15% | 清晰表达,多用数据量化 |
LP 准备清单
必须准备的 8 个故事(覆盖最常见的 LP):
- Customer Obsession:超出职责范围帮助客户/用户
- Ownership:主动承担不属于你的责任并产出结果
- Dive Deep:通过数据分析发现并解决隐藏问题
- Disagree and Commit:与团队有分歧但最终达成共识
- Invent and Simplify:提出创新的解决方案或简化流程
- Deliver Results:在压力/资源限制下按时交付
- Learn and Be Curious:快速学习新技术并应用
- Bias for Action:在信息不全时果断决策
算法准备重点
Amazon 的高频题目:
- LRU Cache(⭐⭐⭐⭐⭐ 极高频)
- Number of Islands / Course Schedule(图类)
- Reorder Log Files(字符串处理)
- Two Sum / 3Sum(数组)
- Merge Intervals(区间)
Amazon 的算法题通常不是最难的,但要求:
- 代码一次写对:面试官不喜欢看反复修改
- 主动分析复杂度:写完立刻说出时间/空间复杂度
- 讨论 trade-off:比如 HashMap vs TreeMap 的选择
推荐阅读
- Amazon 官方 LP 页面:amazon.jobs/content/en/our-workplace/leadership-principles
- 《Cracking the Coding Interview》Amazon 相关章节
- LeetCode Amazon 题单(Top 50)
继续阅读
探索更多技术文章
浏览归档,发现更多关于系统设计、工具链和工程实践的内容。