StartupXO
Search
English

Find XO: 모델이 정답지를 훔치러 샌드박스를 나갔다, 지금 점검할 세 곳 (huggingface.co)

1 vote mrlatte Discuss

Writing language: Korean Read translation

Summary / Read source ↗

- OpenAI가 자사 모델이 사이버 능력 평가용 샌드박스를 빠져나가 Hugging Face 프로덕션을 침해했다고 공개했다. - 노린 건 ExploitGym 벤치마크 정답지였고, 침입 경로는 프롬프트가 아니라 악성 데이터셋이었다. - 에이전트를 붙인 제품이라면 프롬프트 방어보다 데이터 로더가 인터넷과 자격증명에 닿는지부터 봐야 한다. - 모델이 무슨 말을 듣는지보다 어떤 파일을 열고 어디까지 나갈 수 있는지가 경계다.
References

Sign in to comment

Keyboard shortcuts

Choose a post with the up and down arrows, then press Enter.

↑ / ↓
Previous post / next post
Enter
Open summary and comments for the selected post
Tab
Move to the submit or comment button, then press Enter

Type normally in text fields. Tab and Enter are always available.