Tony Wang
Washington, District of Columbia, United States
351 followers
324 connections
View mutual connections with Tony
Tony can introduce you to 10+ people at National Institute of Standards and Technology (NIST)
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
View mutual connections with Tony
or
New to LinkedIn? Join now
By clicking Continue to join or sign in, you agree to LinkedIn’s User Agreement, Privacy Policy, and Cookie Policy.
Activity
351 followers
-
Tony Wang reposted thisTony Wang reposted this🎉 I am extremely pleased to announce that the Center for AI Standards and Innovation (CAISI) has recently released the public draft, Practices for Automated Benchmark Evaluations of Language Models (NIST 800-2): https://lnkd.in/e2XniK5N A 60-day comment period is now open. ‼️ We are requesting public comment on the draft document through March 31, 2026. ‼️ Feedback can be emailed to AI800-2@nist.gov in any form. Useful guidance on AI evaluation is of increasing urgency. The primary audience for NIST AI 800-2 is technical staff at organizations evaluating AI systems, including AI deployers, developers, and third-party evaluators. However, all potential consumers of AI evaluation reports can benefit from robust, well-communicated evaluation practices that advance gold-standard science and inform AI procurement and implementation decisions. The draft organizes practices into three sections and a glossary to clarify and unify the communication of terms: (1) defining evaluation objectives and select benchmarks, (2) implementing and running evaluations, and (3) analyzing and reporting results. Please give it a read. CAISI invites input on any aspect of this draft document, but we are particularly interested in 1. The usefulness and relative importance of the included practices and principles (in general or for specific use cases, types of evaluations, or audiences); 2. The important practices that are within the document’s scope but missing from the draft; Which practices are emerging vs. existing best practices; 3. Any content that is incorrect, unclear, or otherwise problematic; 4. Other common terms in benchmark evaluation that would be useful to define in the glossary, or for the field to use in a standardized manner; 5. When automated benchmark evaluations are more or less useful relative to other evaluation paradigms. See this post for a static link to the above areas: https://lnkd.in/e8C6puXz *Note that all emails, including attachments and other supporting materials, may be subject to public disclosure. Congratulations to Ryan Steed, Drew Keller, Tony Wang, Peter Cihon and the rest of CAISI for this huge achievement!Towards Best Practices for Automated Benchmark EvaluationsTowards Best Practices for Automated Benchmark Evaluations
-
Tony Wang reposted thisTony Wang reposted thisThe National Institute of Standards and Technology (NIST) Center for AI Standards and Innovation (CAISI) is hiring for 𝐀𝐈 𝐄𝐯𝐚𝐥𝐮𝐚𝐭𝐢𝐨𝐧𝐬 𝐒𝐜𝐢𝐞𝐧𝐭𝐢𝐬𝐭, 𝐁𝐢𝐨𝐥𝐨𝐠𝐢𝐜𝐚𝐥 𝐀𝐈 𝐌𝐨𝐝𝐞𝐥𝐬 to advance national security-related evaluations of frontier biological AI models in accordance with the AI Action Plan: https://lnkd.in/ew9HBWwW 𝐏𝐥𝐞𝐚𝐬𝐞 𝐧𝐨𝐭𝐞 that this role is in the "Information Technology Management" job series (https://lnkd.in/epDpEpyQ). The application has strict formatting and content requirements. Applicants 𝐦𝐮𝐬𝐭 submit a resume not to exceed two pages that meets the Basic Requirements, including demonstrating IT-related experience, and all the specialized experience, for the targeted pay band/grade, listed in the posting. Core duties involve: • Designing, developing, and executing national security-related evaluations of frontier biological AI models, such as biological language models and biomolecular modeling and design tools. • Curating relevant benchmarks and ground-truth datasets. • Analyzing biosecurity/national security implications of demonstrated model capabilities. • Collaborating with cross-functional teams of biology, machine learning, and policy subject matter experts. The posting will close on Monday, February 23, 2026.
-
Tony Wang reposted thisTony Wang reposted this‼️ Postdoc opportunity at CAISI – applications due February 1. Please apply and help us get the word out ‼️ The NRC Research Associateship Programs are prestigious postdoctoral and senior research awards designed to provide promising scientists and engineers with high-quality research opportunities at federal laboratories and affiliated institutions. These programs offer a comprehensive experience, including mentorship and access to state-of-the-art facilities, all geared toward enhancing the research career development of the Research Associates. For this opportunity at the Center for AI Standards and Innovation (CAISI) at NIST, proposals are welcome within the scope of measurement and methods for AI evaluations in application, including but not limited to the research areas enumerated here: https://lnkd.in/eRPCk6qF. For example, research to develop and advance metrology and methodologies for evaluating AI systems, including executing evaluations in real settings (e.g., AI evaluation for critical infrastructure, field studies in high risk sectors, longitudinal continuous monitoring of AI post-adoption, etc). Come join our small, but high-powered team! 🧭 Read the full research opportunity listing, and apply to work with the Applied team at CAISI here: https://lnkd.in/eSgHqEM4 Application process described here: https://lnkd.in/e-ziAz4s More info: The Applied Systems team at the Center for AI Standards and Innovation leverages multidisciplinary methodologies to assess, evaluate, and measure AI systems in application and real-world settings to accelerate trustworthy innovation, support the adoption of reliable AI, and protect national security. Our work has wide breadth and reach, including deliverables such as best-practice guidelines, tooling to support evaluations, innovative AI measurement, research papers for publication, workshops and convenings. A successful postdoc applicant would have the opportunity to lead research workstreams, write papers, and engage with AI and measurement science experts from across CAISI and NIST. The artificial intelligence interest group at NIST consists of over 100 scientists who interact regularly via meetings, seminars and informal gatherings. Furthermore, CAISI often collaborates on research with external experts via Cooperative Research and Development Agreements and the AI Consortium. The selected candidate will have the opportunity to contribute to impactful tooling, guidance, research, and best practices that bolster state-of-the-art AI evaluation practices in the US government and the wider field of practitioners.
-
Tony Wang reposted thisTony Wang reposted thisCAISI is recruiting an intern to support an agent security standards project. Position closes Jan. 15 for a February start. Please help spread the word. CAISI will consider exceptional students with demonstrated experience delivering a report or thesis with minimal oversight. Knowledge of AI and/or cybersecurity is a must; familiarity with evaluating AI agent systems and agent-specific security challenges is strongly preferred. The internship runs February-April. It is unpaid; possible full or part time. CAISI has offices in Washington, DC and San Francisco, and may consider remote candidates. Current students (at least half-time in an accredited school) with U.S. citizenship are welcome to apply. To be considered, please request a faculty member provide a paragraph of recommendation in email to peter.cihon@nist.gov no later than January 15. CAISI serves as AI developers’ primary point of contact in USG for security and public-safety evaluations of frontier AI. The internship is a rare opportunity for students to contribute to our mission. For more on CAISI, see https://www.nist.gov/caisi
-
Tony Wang shared thisHi friends, the US AI Safety Institute is looking to expand its technical staff. My experience at the institute so far is that we have a lot of freedom to work on the things that we think are the most impactful. We're looking for folks who can help us take advantage of this opportunity to lay the foundation for the US government's efforts in AI safety. We're pretty small right now, so if you join you'd have an outsized impact. If this speaks to you, I encourage you to apply. Remote: Closes on 7/29/2024; https://lnkd.in/e6vX-p29 DC / CA: Closes on 8/6/2024; https://lnkd.in/e9MTYWnyUSAJOBS connects job seekers with federal jobs across the United States and around the world as the official employment site for the federal governmentUSAJOBS connects job seekers with federal jobs across the United States and around the world as the official employment site for the federal government
-
Tony Wang shared thisExcited to share a new paper I've been working on for the past year! My personal takeaway is that we need to move beyond basic adversarial training if we want to build robust AI agents. By "basic adversarial training", I mean schemes where real examples of failures are collected and then agents are trained to downweight the probability of the actions that led to those failures. I think the issue with basic adversarial training is twofold. Firstly, too much compute needs to be expended for each counterexample (our strongest defense spent 18x more compute on finding attacks compared to learning to defend). Secondly, existing learning algorithms aren't able to extract enough juice from counterexamples (i.e. they don’t have good sample efficiency / generalization). We comment on some potential solutions in our paper: https://lnkd.in/eXkX6Ts8Tony Wang shared this🛡 Is AI robustness possible, or are adversarial attacks unavoidable? We investigate this in Go, testing three defenses to make superhuman Go AIs robust. Our defenses manage to protect against known threats, but unfortunately new adversaries bypass them, sometimes using qualitatively new attacks! 😈 Last year we found that superhuman Go AIs are vulnerable to “cyclic attacks”. This adversarial strategy was discovered by AI, but can be replicated by human players. See our previous update: https://buff.ly/4cqdVYW We were curious to know whether it was possible to defend against the cyclic attack. Over the course of a year, we tested three different ways of patching the cyclic-vulnerability in KataGo, the leading open-source Go AI: 📚 Defense #1: Positional Adversarial Training. The KataGo developers added manually curated adversarial examples to KataGo’s training data. While this successfully defends KataGo against our original versions of the cyclic attack, we find new variants of the cyclic attack that still get through. We also find brand new attacks that defeat this system, such as the “gift attack” shown at https://buff.ly/3xkqKoJ 🎁 🔄 Defense #2: Iterated Adversarial Training. This approach alternates between defense & offense, mirroring a cybersecurity arms race. Each iteration improves KataGo's defense against known adversaries, but after 9 cycles, the most defended model can still be beaten 81% of the time by a novel variant of the cyclic attack we call the “atari attack”: https://buff.ly/3RxW9uI 🎋 🖼️ Defense #3: Vision Transformer (ViT). In this defense, we replaced KataGo’s convolutional neural network (CNN) backbone, which focuses on local patterns, with a ViT backbone, which can attend to the entire board at once. Unfortunately our ViT bot remained vulnerable to the original cyclical attack. Three diverse defenses all being overcome by new attacks is further evidence that AI robustness issues like jailbreaks are likely to remain a problem for many years to come. 💡 However, we did notice one positive sign: defending against any fixed static attack was quick and easy. We think it might be possible to leverage this property to build a working defense both in Go and other settings. In particular, one could a) grow the adversarial training dataset by scaling up attack generation, b) improve the sample efficiency / generalization of adversarial training, and / or c) apply adversarial training online to defend against adversaries as they are learning to attack. For more information: 🔗 Visit our website: https://goattack.far.ai/ 📝 Check out the blog post: https://lnkd.in/eCFYYupX 📄 Read the full paper: https://lnkd.in/eZrSpCc6 👥 Research by Tom Tseng, Euan McLean, Kellin Pelrine, Tony Wang and Adam Gleave. 🚀 If you're interested in making AI systems more robust, we're hiring! Check out our roles at https://far.ai/jobs
-
Tony Wang reacted on thisStanford Institute for Human-Centered Artificial Intelligence (HAI)
Stanford Institute for Human-Centered Artificial Intelligence (HAI)
1wTony Wang reacted on thisCongratulations to Rob Reich, HAI associate director and senior fellow, on being named to the expert group advising Gov. Gavin Newsom on California's AI executive order. The group will develop recommendations on some of the hardest open questions in AI governance, from embedding independent verification organizations inside frontier labs, to verifying the safety frameworks companies file with the state and advancing the creation of a "kill switch" for frontier models. Rob joins Jason Goldman (Center for Shared AI Prosperity), Gillian K. Hadfield (Johns Hopkins, Vector Institute), and Alondra Nelson (Institute for Advanced Study). This role builds on Rob’s previous work as a senior advisor to the U.S. AI Safety Institute: accelerating the scientific practice of AI testing and evaluation so that oversight rests on evidence rather than metaphor. In his words, "Innovation and safety go hand in hand in any responsible marketplace, and AI is no different." Congratulations, Rob! Well deserved. https://lnkd.in/gbWZm8_k -
Tony Wang liked thisComputer Science at the University of Illinois Chicago
Computer Science at the University of Illinois Chicago
3wTony Wang liked thisJonathan Liu joined the Department of Computer Science this fall as a clinical assistant professor. His research focuses on problem-solving skill development in computer science. https://lnkd.in/g2iNtARk -
Tony Wang reacted on thisTony Wang reacted on thisExcited to announce that I joined US CAISI this week as the Lead Research Engineer on their Chem Bio Evals Team! I will be conducting pre-release safety evaluations for frontier AI models on behalf of the federal government 🔎 We have an all star team here at the Center for AI Standards & Innovation in National Institute of Standards and Technology (NIST). If you are a capable engineer with an interest in helping to reduce risks from AI systems, please reach out! We are still hiring 📈 #ai #governance #biosecurity
-
Tony Wang liked thisTony Wang liked thisI've joined California state government! As AI Science Advisor at the Office of Emergency Services (Cal OES), I'm advising senior leadership on frontier AI safety, with a particular emphasis on critical safety incidents, AI and cyber defense, and risk from developers’ internal deployment of AI, such as sabotage by AI agents and automated AI R&D. AI agents are learning to autonomously execute exponentially more complex projects over time. That now includes cyberattacks and will pose risks for critical infrastructure. At Cal OES, I look forward to working with frontier developers in reviewing assessments of risk from their internal use of frontier models. I'm also excited to help California prepare thoughtful, proactive response plans and playbooks for frontier AI risks. I report to Commander of the California Cybersecurity Integration Center (Cal-CSIC) Matthew Sage and work closely with the Homeland Security Cyber Policy Team. I'm placed through the new AI Science Residency Program by the California Council on Science and Technology, a nonpartisan, nonprofit organization established via a unanimous vote of the California Legislature in 1988. https://lnkd.in/gXX3EXXyCCST Launches California AI Science Residency Program on Frontier AI Safety with Appointment of First AI Science Advisors to the California Governor's Office of Emergency Services and Department of Technology - California Council on Science & Technology (CCST)CCST Launches California AI Science Residency Program on Frontier AI Safety with Appointment of First AI Science Advisors to the California Governor's Office of Emergency Services and Department of Technology - California Council on Science & Technology (CCST)
-
Tony Wang reacted on thisTony Wang reacted on thisAfter 19 months of service, today marked my last day with CAISI and the US Government. It was truly a privilege to serve alongside some of the most brilliant and dedicated people I've ever met. When I entered the government, I was driven by the question "could AI ever be good at hacking?" which demanded rigorous metrology, high-fidelity evaluations, and detailed case studies to understand what was possible and what might be coming next. In a very short period of time, the world has rapidly changed, and now I think I know the answer. For me, this means it's time to focus on the next question: what we should do about these AI hacking skills. I want to extend a heartfelt thanks to my CAISI colleagues, interagency partners, industry partners, and everyone who ever joined a TRAINS meeting or sat through one of my cyber briefings. I'm proud of all the things we accomplished together. In a few weeks I'll share where I'm headed next, but first — it's time for a vacation.
-
Tony Wang liked thisTony Wang liked thisMeasuring someone's productivity by their token usage is a horrible idea. Giving everyone the same fixed token budget isn't much better. So what's the right way to roll out AI across your org? We built a system to measure how many productive engineering hours every Devin task is worth, validated against a dataset of real engineers’ times estimates. The goal is to answer the fundamental question that companies are grappling with: how much real value are you getting from each of your agent sessions? On top of that, we're giving an AI productivity guarantee! Now if Devin delivers less engineering value than you're paying for, we fund your usage until it does, up to $10M. The whole industry needs to move from measuring activity to measuring output. We hope to see more AI companies taking this approach.
-
Tony Wang liked thisTony Wang liked thisGo to bed. Same time every night. Non-negotiable. If kids, tell them they’re on their own. You have a schedule to keep. No kids, no excuses. Best thing you can do for yourself. And others. Better mood. More willpower. Clearer mind. Better human.
-
Tony Wang liked thisTony Wang liked thisI'm honored to be one of the few Americans chosen for the AI Scientific Panel. I'm excited to contribute technical expertise here and help make sure U.S. perspectives are represented. AI policy for the most capable models can be more thoughtful when there's pragmatic, independent analysis to inform it.
Experience
Education
Honors & Awards
-
USACO Finalist
United States of America Computing Olympiad
Selected as one of 24 finalists in the annual USACO program.
Organizations
-
Peninsula Youth Orchestra
Violinist, Assistant Concertmaster
-
View Tony’s full profile
-
See who you know in common
-
Get introduced
-
Contact Tony directly
Other similar profiles
Explore more posts
-
University of Cambridge Department of Computer Science and Technology
6K followers
AI researchers are fascinated by Large Language Models. But Large Tabular Models? Nope, they don't get anything like as much research love. Our PhD student Xiangjian Jiang wants to change that. He has just been awarded a prestigious 2025 Global Google PhD Fellowship and hopes the accolade will help him highlight an under-explored area of AI research that has applications in realms from finance to health. https://lnkd.in/gryW5wVc
59
-
Jeffrey Wong
Airbnb • 3K followers
If you are trained in AB testing and linear models, how do you move into deep learning? Linear models are such an important foundation in data science, and especially for AB testing. They power variance reduction, heterogeneous effects, and control for confounders in cases when randomization broke. Inference from a linear model is largely a series of matrix multiplications. Deep learning is too. It's a composition of linear functions stacked with nonlinear activations, and can be thought of as a way to form nonlinear basis functions for regression. These days it is more important than ever to truly understand what your model is doing, so I've written an educational piece on how to pivot a mental model around regression to a mental model around neural networks. Check it out here https://lnkd.in/g24Wr_kw
73
1 Comment -
Princeton Materials Institute
5K followers
In a major step toward practical quantum computers, Princeton engineers have built a superconducting qubit that lasts three times longer than today’s best versions. “The real challenge, the thing that stops us from having useful quantum computers today, is that you build a qubit and the information just doesn’t last very long,” said PMI's Andrew Houck, leader of a federally funded national quantum research center, Princeton’s dean of engineering and co-principal investigator on the paper. “This is the next big jump forward.” In a Nov. 5 article in the journal Nature, the Princeton team reported their new qubit lasts for over 1 millisecond. This is three times longer than the best ever reported in a lab setting, and nearly 15 times longer than the industry standard for large-scale processors. The researchers built a fully functioning quantum chip based on this qubit to validate its performance, clearing one of the key obstacles to efficient error correction and scalability for industrial systems. Read more: https://lnkd.in/e3EGtPeE
4
-
RISE Robotics
989 followers
🤖 The continuum robotics challenge just got solved. Researchers have cracked a unified mathematical model for tendon-actuated concentric tube robots—a breakthrough that combines the best of two worlds. For years, engineers faced a trade-off: tendon-driven robots lacked flexibility in their degrees of freedom, while concentric tube designs suffered from instability. This new Cosserat rod-based framework handles multiple tubes with multiple tendons actuating each one, validated through experiments achieving <4% tip prediction error. The model is so robust it accurately predicts behavior across existing robots in the field. • Why it matters: Accurate modeling means better shape estimation and control—unlocking precision surgical applications, delicate manipulation tasks, and next-gen minimally invasive procedures. Follow for more breakthroughs in robotics and AI innovation! 🚀 #Robotics #ContinuumRobotics #Biomimetics #ControlSystems #RoboticSurgery #Innovation Source: https://lnkd.in/gQtDFW7Z
2
-
Sandia National Laboratories
176K followers
Quantum magic 🪄 Sandia researchers and partners at Quantinuum and University of California, Davis validated a new algorithm for error-resistant quantum computation. In a paper published by American Physical Society, the team details a new and more efficient method for preparing “magic states,” enabling future quantum computers to run programs even if one component fails, a concept called fault-tolerance. “A magic state is a particular superposition state into which a qubit can be initialized,” explained Sandia researcher Robin Blume-Kohout. “It can be difficult to prepare a magic state, but once accomplished, it becomes very useful in calculations. It’s like a $100 gift card that you can redeem for cool stuff whenever you want. Specifically, the quantum computer’s programmer can ‘redeem’ it for a special logic operation, the T gate, that is necessary from time to time in any useful quantum program.” Read more: https://bit.ly/4rEaGWf
97
3 Comments -
Alejandro Ribeiro
University of Pennsylvania • 2K followers
I returned to the #UniversityOfMinnesota on Thursday to talk about Spectral Properties of Graph Neural Networks (GNNs). The message here is that graph convolutions have trouble handling high frequencies, defined as signals that change quickly over the graph. This observation has implications on the stability of GNN outputs to graph deformations and the transferability of GNNs across scales. Stability and transferability properties of GNNs are analogous to stability and transferability properties of CNNs. This is because GNNs and CNNs are algebraically equivalent objects. (See references below). A recording of this talk delivered at the #SimonsFoundation: https://lnkd.in/e4Rxw-gX My graph neural networks course at #Penn: https://gnn.seas.upenn.edu Stability properties of GNNs: https://lnkd.in/en-UH8jj Transferability properties of GNNs: algehttps://lnkd.in/efrnDJFM Algebraic neural networks: manifold https://lnkd.in/eDpxp-tj Manifold convolutions: https://lnkd.in/eraqxeDP Learning by transference: https://lnkd.in/etaPVQVS Learnable perception action communication loops: chttps://lnkd.in/ejZ5qKkC Covariance Neural Networks: https://lnkd.in/ee2CBweA For those of you that have never lived in #Minnesota, it is always a treat to return to the best state in the union.
143
1 Comment -
Candace Gillhoolley
Data Driven Media, LLC • 20K followers
Imagine quantum key distribution (QKD) systems so small they fit on a single chip. That's the future. Researchers are already dedicating years to this, creating compact, efficient systems. Soon, this technology will be as ubiquitous as your GPU or CPU. While the exact applications and costs are still evolving, one thing is clear: quantum tech is on the verge of a major breakthrough. #QuantumTech #QKD #Cybersecurity #Innovation
4
-
The Real Preneur
1K followers
LLNL’s HPC Legacy: Judy Hill on Innovations and Challenges Lawrence Livermore National Laboratory (LLNL) leads the charge in high-performance computing (HPC). Judy Hill, Deputy for High Performance Computing, offers insights into Livermore Computing. This center drives groundbreaking simulations that fuel discovery and innovation. HPC now ranks high on U.S. priorities. It saves energy, cuts emissions, boosts competitiveness, and cements America's tech leadership. At Department of Energy (DOE) sites like LLNL, HPC stands as the third pillar alongside theory and experiment....
-
Princeton ECE
5K followers
A new feature from Nature titled "Quantum computers will finally be useful: what’s behind the revolution" highlights research from Princeton ECE professors Andrew Houck and Nathalie de Leon. The piece describes how their work improving the performance of qubits has contributed to dramatically shortening the projected timeline for practical quantum computers. "Just a few years ago, many researchers in quantum computing thought it would take several decades to develop machines that could solve complex tasks, such as predicting how chemicals react or cracking encrypted text," according to the article. "But now, there is growing hope that such machines could arrive in the next ten years." Check it out 👉 https://lnkd.in/eC5cSDs4
19
1 Comment -
DIGZON
834 followers
[PDF] Thermodynamics for the practicing engineer Theodore L., Ricci F., Vliet T.V. https://lnkd.in/eTUW5TGP This book concentrates specifically on the applications of thermodynamics, rather than the theory. It addresses both technical and pragmatic problems in the field, and covers such topics as enthalpy effects, equilibrium thermodynamics, non-ideal thermodynamics and energy conversion applications. Providing the reader with a working knowledge of the principles of thermodynamics, as well as experience in their application, it stands alone as an easy-to-follow self-teaching aid to practical applications and contains worked examples. digzon #simple #Engineering #RicciF. #TheodoreL. #VlietT.V. https://lnkd.in/ePB9Apvd
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content